- Research Article
57
- 10.1016/j.icte.2020.05.003
Implementing action mask in proximal policy optimization (PPO) algorithm
- May 20, 2020
- ICT Express
- Cheng-Yen Tang + 3 more +3
Implementing action mask in proximal policy optimization (PPO) algorithm
Industry 5.0 promotes the transformation of manufacturing toward flexibility, personalization, and sustainability. As a critical component of closed-loop manufacturing systems, disassembly operations urgently require more flexible and efficient human–robot collaboration models. To this end, this work, for the first time, proposes a multihuman–robot collaborative circular disassembly line balancing problem. By allowing workers to move between robotic workstations, the proposed system enhances operational flexibility. Furthermore, a multiworker mechanism is introduced to improve fault tolerance and system stability, overcoming the limitations of fixed worker positions in existing collaborative disassembly research. To solve this problem, we formulate a discrete-time mixed-integer programming model based on product AND/OR graphs, aiming to maximize disassembly profit. The model’s correctness is verified using CPLEX. Additionally, we develop a heterogeneous graph neural network-enhanced proximal policy optimization (PPO) algorithm. By integrating product and workstation information into a heterogeneous graph, the algorithm performs two-stage feature extraction and node embedding via graph neural networks. Based on these embeddings, the agent dynamically selects multiple actions per decision step to simulate the behavior of multiple workers moving simultaneously. Experimental results show that the proposed method outperforms traditional reinforcement learning algorithms such as PPO and deep Q-network algorithm in terms of disassembly profit. Moreover, it demonstrates strong generalization capability in cross-task transfer and scalability experiments involving different task graph sizes. The improved performance is achieved with acceptable computational time.
Implementing action mask in proximal policy optimization (PPO) algorithm
Implementing action mask in proximal policy optimization (PPO) algorithm
Multiple-UAV Reinforcement Learning Algorithm Based on Improved PPO in Ray Framework
Distributed multi-agent collaborative decision-making technology is the key to general artificial intelligence. This paper takes the self-developed Unity3D collaborative combat environment as the test scenario, setting a task that requires heterogeneous unmanned aerial vehicles (UAVs) to perform a distributed decision-making and complete cooperation task. Aiming at the problem of the traditional proximal policy optimization (PPO) algorithm’s poor performance in the field of complex multi-agent collaboration scenarios based on the distributed training framework Ray, the Critic network in the PPO algorithm is improved to learn a centralized value function, and the muti-agent proximal policy optimization (MAPPO) algorithm is proposed. At the same time, the inheritance training method based on course learning is adopted to improve the generalization performance of the algorithm. In the experiment, MAPPO can obtain the highest average accumulate reward compared with other algorithms and can complete the task goal with the fewest steps after convergence, which fully demonstrates that the MAPPO algorithm outperforms the state-of-the-art.
Read moreProximal policy optimization via enhanced exploration efficiency
Proximal policy optimization via enhanced exploration efficiency
Research on Helicopter Flight Trajectory Control Based on Proximal Policy Optimization Algorithm
Currently, helicopter towing operations mainly rely on towing systems and multi-person coordinated control, resulting in low work efficiency and posing a threat to life safety. With the development of artificial intelligence and computer science, automating helicopter towing systems is key to achieving safe and efficient helicopter operations. This paper proposes a helicopter tracking control method based on the Proximal Policy Optimization (PPO) algorithm for the helicopter towing trajectory tracking problem. This paper first improves the PPO algorithm network structure and sets the hyperparameters of the algorithm. Then, a deep learning environment is constructed based on the Markov decision process, and the state space and action space are reasonably set. A reward function is designed based on lateral deviation, longitudinal deviation, and heading deviation. Finally, the PPO algorithm is used to train the trajectory tracking control strategy, and the optimal tracking control strategy is obtained and verified by simulation. The simulation results show that the helicopter tracking control strategy based on the PPO algorithm has a fast convergence speed, and the tracking effect verifies the feasibility and generalization of the algorithm.
Read moreDeep reinforcement learning-based digital twin for droplet microfluidics control
This study applied deep reinforcement learning (DRL) with the Proximal Policy Optimization (PPO) algorithm within a two-dimensional computational fluid dynamics (CFD) model to achieve closed-loop control in microfluidics. The objective was to achieve the desired droplet size with minimal variability in a microfluidic capillary flow-focusing device. An artificial neural network was utilized to map sensing signals (flow pressure and droplet size) to control actions (continuous phase inlet pressure). To validate the numerical model, simulation results were compared with experimental data, which demonstrated a good agreement with errors below 11%. The PPO algorithm effectively controlled droplet size across various targets (50, 60, 70, and 80 μm) with different levels of precision. The optimized DRL + CFD framework successfully achieved droplet size control within a coefficient of variation (CV%) below 5% for all targets, outperforming the case without control. Furthermore, the adaptability of the PPO agent to external disturbances was extensively evaluated. By subjecting the system to sinusoidal mechanical vibrations with frequencies ranging from 10 Hz to 10 KHz and amplitudes between 50 and 500 Pa, the PPO algorithm demonstrated efficacy in handling disturbances within limits, highlighting its robustness. Overall, this study showcased the implementation of the DRL+CFD framework for designing and investigating novel control algorithms, advancing the field of droplet microfluidics control research.
Read moreProximal Policy Optimization-Based Hierarchical Decision-Making Mechanism for Resource Allocation Optimization in UAV Networks
To address the resource allocation problem in dynamic environments where multiple unmanned aerial vehicle base stations (UAV-BSs) provide efficient downlink services to ground users, this paper proposes a novel hierarchical decision-making mechanism based on the Proximal Policy Optimization (PPO) algorithm. The proposed method optimizes time-frequency resource allocation in the downlink, aiming to maximize the total user throughput over multiple time slots. By constructing channel and interference models, the complex multi-channel resource allocation problem is decomposed into a series of single-channel decision subproblems, significantly reducing the action space complexity. Specifically, the original exponential complexity O(NM) (where N is the number of users and M is the number of channels) is reduced to a linear complexity O(N), effectively alleviating the curse of dimensionality. Simulation results demonstrate that the proposed hierarchical architecture, integrated with the PPO algorithm, achieves superior performance in terms of total throughput, convergence speed, and stability compared to existing methods. This study provides new insights and technical support for efficient resource management in UAV-BS systems operating in complex and dynamic environments.
Read moreReinforcement learning-based control with application to the once-through steam generator system
Reinforcement learning-based control with application to the once-through steam generator system
Parameter Identification Method of Grid-Forming Static Var Generator Based on Trajectory Sensitivity and Proximal Policy Optimization Algorithm
As the penetration rate of new energy continues to increase, the active voltage support capability of the power system is decreasing. The grid-forming static var generator (GFM-SVG) features the advantages of fast dynamic response, strong reactive power support, and high overload capacity, which play an important role in maintaining voltage stability. However, the parameters of the GFM-SVG are often unknown due to trade secret reasons. Meanwhile, the parameters may be changed during the long-term operation of the system, which brings challenges to the system stability analysis and control. Aiming at this problem, a parameter identification method based on trajectory sensitivity analysis and the proximal policy optimization (PPO) algorithm is proposed in this paper. Firstly, through trajectory sensitivity analysis, the key influential parameters on the output characteristics of the GFM-SVG can be selected, which can reduce the dimensionality of the identification parameters and improve the identification efficiency. Then, a parameter identification framework based on the PPO algorithm is constructed for GFM-SVGs, which utilizes its adaptive learning capability to achieve accurate identification of the key parameters of the system. Finally, the effectiveness of the proposed parameter identification method is verified through simulation examples. The simulation results show that the identification error of the parameters in the GFM-SVG is small. The proposed method can characterize the output response of the GFM-SVG under different operating conditions.
Read moreResearch on scheduling optimization method of pumped storage system based on PPO and multi-objective optimization
With the rapid development of renewable energy, power systems face increased flexibility and dispatch requirements, particularly in response to grid load fluctuations and energy reserve security. Pumped hydro storage, a mature energy storage technology, is widely used in power regulation and reserve capacity provision, but its dispatching still presents significant optimization challenges. Traditional dispatching methods struggle to simultaneously meet multi-objective optimization requirements. Therefore, this study proposes a dispatching optimization method based on Proximal Policy Optimization (PPO) combined with Multi-Objective Optimization (MOO) to maximize profit, ensure sufficient reserve capacity, and safeguard grid stability. We construct a multi-objective optimization model for pumped hydro storage and combine it with the Proximal Policy Optimization (PPO) algorithm for dynamic dispatch optimization. A multi-objective reward function is used to weightedly optimize profit, reserve capacity, and grid stability. Experimental results demonstrate that the PPO + MOO method significantly outperforms traditional methods (such as piecewise linear approximation, genetic algorithm, and particle swarm optimization) across multiple evaluation metrics. It effectively improves dispatching accuracy and system stability, particularly in complex power markets and under conditions of large load fluctuations. Furthermore, the PPO algorithm demonstrates excellent convergence and stability during training, ensuring continuous model optimization in dynamic environments. Although computational time is relatively long, optimization techniques such as weight pruning have significantly improved computational efficiency. Ultimately, this study provides a new solution for the scheduling optimization of pumped hydro energy storage systems and provides an important reference for scheduling problems in smart grids and renewable energy systems.
Read moreIntelligent Traffic Signal Control with Deep Reinforcement Learning at Single Intersection
In this paper, we apply the Proximal Policy Optimization (PPO) algorithm in intelligent traffic signal control at a single intersection with eight lanes and four signal phases. The optimization goal is to minimize the average waiting time of vehicles so as to improve the traffic efficiency of the intersection. Extensive experiments are conducted in Simulation of Urban MObility (SUMO) to evaluate the performance of the proposed algorithm, and compare it with other classic algorithms including Deep Q-network (DQN), Advantage Actor Critic (A2C) and Fixed Time. Simulation results show that the proposed PPO algorithm outperforms the others under various traffic scenarios to different extent. The performance gain is significant under unbalanced traffic where one direction is saturated while the other is not, and becomes marginal when all the directions are saturated or unsaturated. PPO also demonstrates good portability and robustness over time-varying traffic patterns, while implies it could be a preferable option for implementation in real world intelligent traffic signal control systems.
Read moreA Reinforcement Learning Framework for PAPR Minimization in MU‐MIMO With GNN‐Based CSI Encoding
In Multiuser Multiple‐Input Multiple‐Output systems, the peak‐to‐average power ratio poses a significant constraint, which impacts power efficiency, signal integrity, and overall system reliability. For reducing peak‐to‐average power ratio, although conventional methods provide varying levels of effectiveness, these techniques frequently encounter issues related to requirements for side information, limited flexibility in adapting to changing channel conditions, and computational complexity. To address these challenges, this paper proposes a novel Proximal Policy Optimization (PPO)‐based Graph Neural Network framework for dynamic precoding optimization, particularly focusing on minimizing the peak‐to‐average power ratio in Multiuser Multiple‐Input Multiple‐Output systems. Because of its effective balance between performance and training reliability, the PPO algorithm is employed, which promotes stable and efficient learning. The model uses the CSI data to build a graph, and then extracts features through a GNN to obtain spatial correlations. These extracted GNN features were used as the state of a PPO reinforcement learning agent. The PPO agent is trained to find optimal precoding policies to minimize PAPR without compromising signal quality. Then, the acquired policy is applied to precode at the transmitter that transmits OFDM signals with low PAPR and high spectral efficiency. Extensive MATLAB simulations confirm the efficacy of the proposed framework, demonstrating a notable peak‐to‐average power ratio reduction of 4.8 dB at a complementary cumulative distribution function of 10 −3 . Even in fluctuating channel conditions and varying user densities, the framework achieves improved signal‐to‐interference‐plus‐noise ratio and bit error rate performance, while keeping computational complexity low. The proposed model eliminates the necessity for side information, guaranteeing seamless compatibility with existing Multiuser Multiple‐Input Multiple‐Output transceiver designs and enabling real‐time applications. These findings emphasize the proposed model's potential as an efficient, robust, and scalable method for future wireless communication systems.
Read moreUsing Reinforcement Learning to Identify The Key Factors for Players to Win Games
With the gradual development of sports data analysis, data-driven game prediction has gradually become an important tool for improving the decision-making efficiency of teams and coaches. Basketball, as a team sport, is influenced by multiple factors, and traditional statistical analysis methods make it difficult to intuitively identify the factors that have the greatest impact on the game. This study proposes a reinforcement learning model based on the Proximal Policy Optimization (PPO) algorithm for predicting the winning rate of NBA games. By collecting career statistics of NBA players and combining them with the team's current season winning rate, a feature vector containing individual player characteristics and team winning rate is constructed. The model is trained using the PPO algorithm to identify the most critical features that affect the winning rate of games. This study not only provides new ideas for sports event prediction, but also provides solutions for future decision-making and dataization in big data environments.
Read moreEnergy-efficient Distributed Heterogeneous Hybrid Flow-shop Scheduling Using Graph Neural Network and Deep Reinforcement Learning
With growing environmental awareness and increasing energy demands, sustainable manufacturing has become a focal point in the industry. Meanwhile, globalization has propelled distributed manufacturing systems as a dominant trend. This paper tackles the energy-efficient distributed heterogeneous hybrid flow-shop scheduling problem (EDHHFSP), aiming to minimize both makespan and total energy consumption. We first formulate a mixed-integer linear programming (MILP) model to provide a benchmark for small instances. More importantly, we propose a novel end-to-end deep reinforcement learning framework based on a heterogeneous graph neural network, which models the scheduling problem as a distributed decision-making process. A key innovation lies in the design of an action space composed of "job–factory" and "operation–machine" pairs, enabling fine-grained, decentralized scheduling decisions. Our approach starts with a novel heterogeneous graph representation of scheduling states, capturing complex interactions among jobs, factories, and machines. A three-stage embedding mechanism is developed to encode real-time scheduling environments. The agent then learns a parameterized policy using the proximal policy optimization (PPO) algorithm, guided by a reward function that balances makespan and energy efficiency. Experimental results demonstrate that our method generalizes well across different problem scales and significantly outperforms traditional heuristics and learning-based baselines in terms of both scheduling quality and energy savings.
Read moreReinforcement learning based fractional fuzzy controller for photovoltaic systems
Reinforcement learning based fractional fuzzy controller for photovoltaic systems
Tuning of fuzzy controller with arbitrary triangular input fuzzy sets based on proximal policy optimization for time-delays system
Tuning of fuzzy controller with arbitrary triangular input fuzzy sets based on proximal policy optimization for time-delays system
Read more