- Research Article
- 10.3934/dcdss.2026036
Fixed-wing UAV swarm path finding based on heuristic independent Q-learning with shared buffer
- Jan 01, 2026
- Discrete and Continuous Dynamical Systems - S
- Lingzhi Wang + 3 more +3
This study addresses the challenges encountered by traditional multi-agent pathfinding methods in the applications of fixed-wing UAVs, including inadequate adaptation to dynamic constraints, lack of energy optimization, low exploration efficiency, and inefficient data utilization. To tackle these issues, we propose an improved independent Q-learning algorithm that incorporates the dynamics of fixed-wing UAVs. Initially, we expand the traditional multi-agent pathfinding action space and incorporate constraints to develop a model that is compatible with the motion characteristics of fixed-wing UAVs. Next, we develop a multi-dimensional reward function system based on independent Q-learning. Additionally, we introduce a heuristic policy that utilizes directional weights, effectively balancing exploration efficiency and algorithm convergence speed. Finally, we implement a shared experience replay buffer and a reward reconstruction mechanism to enable multi-agent experience sharing and collaborative optimization. Experimental results indicate that our proposed algorithm surpasses the A* algorithm in total path energy consumption. Furthermore, the heuristic policy demonstrates a twofold improvement in convergence speed and the success rate of path planning compared to the $ \varepsilon $-greedy policy. The shared experience replay buffer also shows enhanced performance in convergence speed.
Read more