Deep reinforcement learning for two-player fighting game based on opponent pool
LIANG Rong-qin
ZHU Yuan-heng
ZHAO Dong-bin
Abstract:In the realm of gaming artificial intelligence,two-player games represent a fundamental and crucial issue,with one-on-one zero-sum fighting games standing as one of the most quintessential forms of two-player games.In this paper,we explore adversarial strategies for fighting games based on deep reinforcement learning.We begin by constructing a model of the fighting game environment,formulating the states,actions,and reward functions that are applicable to decision-making within these games.We then employ phasic policy gradient algorithms for the learning of adversarial strategies.In pursuit of mastering Nash equilibrium strategies to triumph over any opponent,we construct an opponent pool based on intelligent agents from previous competitions for the purpose of training.We also investigate the impact of opponent selection mechanisms on the training process.Lastly,building on a fixed opponent pool,we devise a self-expanding opponent pool algorithm to enhance the comprehensiveness of the opponent strategies and bolster the robustness of the trained agents.To expedite the process of environment sampling,we leverage conventional parallel architectures and create a distributed,multi-server parallel sampling scheme optimized for two-player games.Experimental comparisons reveal that agents trained using the self-expanding opponent pool method achieve a 96.6%win rate against agents in the fixed opponent pool.Furthermore,they also exhibit a 72.2%win rate when pitted against three agents used solely for testing purposes.
Keywords:real-time fighting gamedeep reinforcement learningtwo-player zero-sum gameopponent policy pool
Publication Date:2025-02-28
Online Publishing Date:2025-08-15(First online date of this platform, not the publication date of the document)
Pages:9( 226-234 )
Control Theory & Applications

Control Theory & Applications

ISTICPKUEICSCD
ISSN:1000-8152
Year, Vol.(Issue):2025,42(2)