PID parameter optimization based on TD3 algorithm of double replay buffer
ZHONG Hao-jun
WANG Zhen-lei
Abstract:PID controller is widely used in the field of industrial control,the selection of its parameters is over-dependent on manual experience,the efficiency is low and the process is complicated.In recent years,deep reinforcement learning has been successfully applied in many fields because of its ability to self-learn from complex environments.In this paper,a PID parameter optimization method based on twin delayed deep deterministic policy gradient(TD3)algorithm of double replay buffer is proposed,and the parameters of PID controller are optimized by deep reinforcement learning.In the whole optimization process,the control problem is regarded as a sequence decision process.The optimization process of PID parameters is transformed into the updating process of the weights of the agent's network by designing the state space,action space and the network structure of the agent.At the same time,to solve the problem of low exploration efficiency in the early stage of TD3 algorithm training,the double experience replay buffer mechanism is added on the basis of TD3 algorithm to improve the efficiency of the early stage of algorithm training.Finally,simulations are performed on the second-order system and first order plus delay time system,and compared with the PID parameter optimization method based on particle swarm optimization(PSO)algorithm.The experimental results show that the PID parameters optimized by the proposed algorithm have better control performance than the PSO algorithm.
Keywords:PID parameter optimizationdeep reinforcement learningTD3
Publication Date:2026-01-30
Online Publishing Date:2026-02-05(First online date of this platform, not the publication date of the document)
Pages:10( 139-148 )
Control Theory & Applications

Control Theory & Applications

ISTICPKUEICSCD
ISSN:1000-8152
Year, Vol.(Issue):2026,43(1)