- Research Article
- 10.1049/cth2.70097
Multi‐Agent Reinforcement Learning Algorithm Based on Local Observation Imitation Learning
- Jan 01, 2025
- IET Control Theory & Applications
- Hui Zhang + 3 more +3
This paper investigates the error accumulation problem in centralized training and decentralized execution (CTDE) policy‐based multi‐agent reinforcement learning (MARL) algorithms, which arises from local observation inaccuracies. To address this issue, we propose a novel MARL algorithm that incorporates imitation learning using local observations. Firstly, by analysing the multi‐agent proximal policy optimization algorithm and examining the problems arising when global states are replaced with local observations, it is proved that insufficient observations can lead to information loss, thereby introducing errors of advantage function, and it is demonstrated that the generalized advantage estimation method accumulates errors during the training process. Then, imitation learning is introduced and a novel training framework that combines reinforcement learning and imitation learning is proposed. During the reinforcement learning phase, an MARL agent trained with global observations acts as an expert. Subsequently, imitation learning is applied to train another agent that mimics the expert's decisions using only local observations. Finally, the effectiveness of this algorithm is verified in some commonly used multi‐agent environments, which demonstrates its superior performance compared to traditional multi‐agent reinforcement learning algorithms.
Read more