- Article
- Open access
- Published:
Scientific Reports (2026) Cite this article
We are providing an unedited version of this manuscript to give early access to its findings. Before final publication, the manuscript will undergo further editing. Please note there may be errors present which affect the content, and all legal disclaimers apply.
Abstract
Real-time rescheduling in Flexible Manufacturing Systems (FMS) under equipment fault disturbances has long faced the core contradiction of balancing response speed and scheduling quality. To address this issue, this paper proposes a cooperative policy learning algorithm based on Heterogeneous Graph Neural Network (HGNN) and Proximal Policy Optimization (PPO). The manufacturing system is modeled as a dynamic graph structure with workpieces, machines, and processes as heterogeneous nodes and precedence, attribution, and manufacturability relationships as heterogeneous edges. When a fault is triggered, the corresponding nodes and associated edges are frozen simultaneously to update the graph topology in real time. HGNN aggregates similar features and cross-type heterogeneous information at the node layer and semantic layer, respectively, through a hierarchical heterogeneous attention mechanism to generate a process embedding vector that integrates global manufacturing state and local fault information. This vector is input into the PPO network; the Actor network outputs the process-machine assignment probability distribution; the Critic network estimates the state value; the clipped surrogate objective constrains the policy update step size to prevent policy collapse under extreme fault scenarios. During the training phase, the fault time, duration, and faulty machines are uniformly and randomly sampled. Minimizing the makespan is used as the primary reward signal to drive end-to-end collaborative optimization. During the deployment phase, a single forward inference outputs the complete rescheduling scheme. Experimental results show that on all benchmark datasets of Brandimarte and Hurink, the proposed framework achieves a normalized makespan convergence rate between 103.9% and 105.1% in large-scale scenarios. The average utilization rate of available machines after fault recovery remains stable at 94.0% to 94.5% in medium- to large-scale scenarios. The end-to-end inference time is less than 12 ms under different combinations of scale and fault intensity. The feasibility rate reaches 94.1% in the high fault rate range (λf > 0.50) outside the training distribution, validating the effectiveness of the proposed framework in balancing real-time performance and scheduling quality under fault perturbations.
Subjects
Acknowledgements
Not applicable.
Ethics declarations
Competing interests
The authors declare that they have no conflicts of interest to report regarding the present study.
Additional information
Publisher’s note
Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.
Rights and permissions
Open Access This article is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License, which permits any non-commercial use, sharing, distribution and reproduction in any medium or format, as long as you give appropriate credit to the original author(s) and the source, provide a link to the Creative Commons licence, and indicate if you modified the licensed material. You do not have permission under this licence to share adapted material derived from this article or parts of it. The images or other third party material in this article are included in the article’s Creative Commons licence, unless indicated otherwise in a credit line to the material. If material is not included in the article’s Creative Commons licence and your intended use is not permitted by statutory regulation or exceeds the permitted use, you will need to obtain permission directly from the copyright holder. To view a copy of this licence, visit http://creativecommons.org/licenses/by-nc-nd/4.0/.
Reprints and permissions
About this article
Cite this article
Lyu, R., Meng, Y. & Wang, W. Optimizing real-time rescheduling mechanism for flexible manufacturing systems considering equipment fault disturbances using HGNN-PPO cooperative policy learning algorithm. Sci Rep (2026). https://doi.org/10.1038/s41598-026-65082-7
Download citation
Received:
Accepted:
Published:
DOI: https://doi.org/10.1038/s41598-026-65082-7