Recent progress of deep reinforcement learning:from AlphaGo to AlphaGo Zero
TANG Zhen-tao
SHAO Kun
ZHAO Dong-bin
ZHU Yuan-heng
Abstract:In the early 2016,the defeat of Lee Sedol by AlphaGo became the milestone of artificial intelligence.Since then,deep reinforcement learning(DRL),which is the core technique of AlphaGo,has received widespread attention,and has gained fruitful results in both theory and applications.In the sequel,AlphaGo Zero,a simplified version of AlphaGo, masters the game of Go by self-play without human knowledge.As a result,AlphaGo Zero completely surpasses AlphaGo, and enriches humans'understanding of DRL.DRL combines the advantages of deep learning and reinforcement learning, so it is able to perform well in high-dimensional state-action space, with an end-to-end structure combining perception and decision together.In this paper, we present a survey on the remarkable process made by DRL from AlphaGo to AlphaGo Zero.We first review the main algorithms that contribute to the great success of DRL,including DQN,A3C, policy-gradient,and other algorithms and their extensions.Then,detailed introduction and discussion on AlphaGo Zero are given and its great promotion on artificial intelligence is also analyze.The progress of applications with DRL in such areas as games,robotics,natural language processing,smart driving,intelligent health care,and related resources are also presented.In the end,we discuss the future development of DRL,and the inspiration on other potential areas related to artificial intelligence.
Keywords:deep reinforcement learningAlphaGo Zerodeep learningreinforcement learningartificial intelligence
Publication Date:2017-01-01
Online Publishing Date:2025-08-15(First online date of this platform, not the publication date of the document)
Pages:18( 1529-1546 )
Control Theory & Applications

Control Theory & Applications

PKUISTICEI
ISSN:1000-8152
Year, Vol.(Issue):2017,34(12)