Probably approximately correct reinforcement learning solving continuous-state control problem
ZHU Yuan-heng
ZHAO Dong-bin
Abstract:One important factor of reinforcement learning (RL) algorithms is the online learning time.Conventional algorithms such Q-learning and state-action-reward-state-action (SARSA) can not give the quantitative analysis on the upper bound of the online learning time.In this paper,we employ the idea of probably approximately correct (PAC) and design the data-driven online RL algorithm for continuous-time deterministic systems.This class of algorithms efficiently record online observations and keep in mind the exploration required by online RL.They are capable to learn the nearoptimal policy within a finite time length.Two algorithms are developed,separately based on state discretization and kd-tree technique,which are used to store data and compute online policies.Both algorithms are applied to the two-link manipulator to observe the performance.
Keywords:reinforcement learningprobably approximately correctkd-treetwo-link manipulator
Publication Date:2016-01-01
Online Publishing Date:2025-08-15(First online date of this platform, not the publication date of the document)
Pages:11( 1603-1613 )
Control Theory & Applications

Control Theory & Applications

PKUISTICEI
ISSN:1000-8152
Year, Vol.(Issue):2016,33(12)