An end-to-end deep reinforcement learning method for multi-order dynamic flexible job shop scheduling problem
WANG Xu
LI Huan
HAN Yuyan
WANG Yuting
WANG Yakun
Abstract:To address the Dynamic Flexible Job Shop Scheduling Problem with Order Random Arrival(DFJSP_ORA),a modeling and solution framework tailored for the actual production environment is pro-posed.First,a mathematical model is formulated to minimize the maximum completion time.The fluid model is then introduced to continuously approximate system behavior and extract key state features.Sub-sequently,the scheduling process is modeled as a Markov Decision Process(MDP),and a deep reinforce-ment learning method based on Proximal Policy Optimization(PPO)is developed to solve the problem.It combines the discrete action space driven by composite rules and the strategy optimization mechanism driv-en by the advantage function,achieving efficient decision-making in dynamic environments.Experimental results demonstrate that the proposed approach performs well in dynamic scheduling scenarios and effec-tively handles uncertainty and complexity in production,providing an efficient and flexible solution for DFJSP_ORA.
Keywords:flexible job shop schedulingdeep reinforcement learningproximal policy optimizationfluid modelmaximum completion time
Publication Date:2026-04-25
Online Publishing Date:2026-08-28(First online date of this platform, not the publication date of the document)
Pages:14( 192-204,273 )
