Convergence Analysis of Multistep Reinforcement Learning Algorithm
YANG Rui
Abstract:Recently,a new algorithm called Q(σ) has been presented to evalued value function in the theory of reinforcement learning algorithm,where σ is the degree of sampling. Q(σ) is a new method between full-sampling and no-sampling and it unifies Sarsa and Expected Sarsa. However,the original paper only tests the performance of Q(σ) on experiments. This paper gives a theo?retical analysis of Q(σ) . It gives a proof that under some conditions,Q(σ) can converge to the value functions.
Keywords:reinforcement learningvalue function estimateoptimizationtemporal difference
Publication Date:2019-01-01
Online Publishing Date:2025-08-15(First online date of this platform, not the publication date of the document)
Pages:4( 1582-1585 )
