• Home
  • Search
  • Decision-making with cooperative multi-agent reinforcement learning
  • https://doi.org/10.32657/10356/183129Copy DOI Icon

Decision-making with cooperative multi-agent reinforcement learning

  • Jan 1, 2025
  • Jing Sun
Show More
  • Abstract
  • Literature Map
  • Similar Papers
Abstract

Reinforcement learning, a machine learning technique, has already shown advancement in solving complex sequential decision-making problems. Many tasks involve multiple agents and require sequential decision-making policies to achieve common goals, such as warehouse automation, autonomous driving, and game-playing. To derive the policies for all agents, these problems can be modelled as multi-agent systems and addressed by multi-agent reinforcement learning (MARL). However, optimizing policies in multi-agent scenarios presents significant challenges due to the intricate behaviours of multiple agents and the non-stationary nature of the environments’ complex dynamics. Firstly, optimizing policies in multi-agent scenarios presents significant challenges due to the complexity of multi-agent behaviours, especially in partially observable environments. Also, the dynamic nature of agents’ behaviours and their interactions with other agents lead to changes in the environment’s states and agents’ observations over time, which is more complicated in open environments. Furthermore, the need to balance individual and collective objectives in certain real-world multi-agent environments also complicates the decision-making process. This doctoral thesis aims to tackle the following three fundamental multi-agent research problems and propose a solution for each. Our research covers from theoretical analysis to practical application. We start by studying the problem of learning an efficient policy under partially observable environments. We notice that a group of agents cooperates to com- pete against another group of agents (named opponents), with the limitation that information about the opponents is inaccessible due to partial observability. To address this issue, we propose a novel multi-agent distributional actor-critic algorithm to achieve speculative opponent modeling with purely local information (i.e., the controlled agent’s observations, actions, and rewards). The actor maintains a speculated belief of the opponents, which we call the speculative opponent models, to predict opponents’ actions using local observations and make decisions accordingly. The distributional critic models the return distribution of the policy. It reflects the quality of the actor and thus can guide the training of the imaginary opponent model that the actor relies on. Extensive experiments confirm that our method successfully models opponents’ behaviours without their data and delivers superior performance against baseline methods with a faster convergence speed. Furthermore, in some environments, teammates’ numbers and policies change with the market, requiring workers to adapt to perform different sets of tasks across time. To solve this issue, we propose an RL-based method that allows the controlled agent to collaborate with dynamic teammates in open environments. The controlled agent maintains a dual teamwork situation inference model to capture the current teamwork state and facilitate reasonable decision-making under partial observability. Considering the dynamic types of teammates, we first leverage the Chinese Restaurant Process-based model to categorize versatile teammate policies into distinct clusters, which improves the efficiency of identifying current teamwork situations. Next, to model the heterogeneous relationships among agents and accommodate the varying number of teammates, we apply heterogeneous graph attention neural networks to learn the representation of the teamwork situation. Extensive experiments confirm that our method outperforms state-of-the-art baselines with faster convergence in various ad hoc teamwork tasks. Lastly, in certain real-world applications, such as routing problems and warehouse management, decision-makers must balance overall benefits with individual fairness among agents. Achieving both learning efficiency and fairness concurrently presents a complex, multi-objective, joint-policy optimization challenge. Moreover, existing approaches are predominantly confined to simulation environments. To address the above issues, we present a pioneering MARL approach to balance individual and collective objectives for agents collaboration. Extensive experiments on synthetic and real-world datasets demonstrate that our method not only surpasses the state-of-the-art DRL methods, but also exhibits advantages over established heuristics, notably in achieving significantly faster optimization speeds. This method underlines the essential integration of fairness into real-world applications, representing a major advancement in creating equitable and efficient logistics solutions. To conclude, this doctoral thesis investigates three fundamental multi-agent decision-making research problems that are ubiquitous and unsolved. The proposed three MARL methodology solutions achieve efficient policy training and performance for agents in multi-agent environments with uncertainties raised by the partially observable, the open environment, and the individual-collective objectives of MARL. This thesis delves into novel designs of various critical components of MARL, including MDP formulations, policy networks, training algorithms, and inference methods. These contributions significantly elevate the effectiveness and efficiency of cooperative MARL, establishing new performance benchmarks.

Similar Papers
  • Research Article
  • Citations4

Multiagent Continual Coordination via Progressive Task Contextualization.

  • Apr 01, 2025
  • IEEE transactions on neural networks and learning systems
  • Lei Yuan +5
  • Research Article

Learning Optimal Policies With Local Observations for Cooperative Multiagent Reinforcement Learning.

  • Mar 20, 2026
  • IEEE transactions on neural networks and learning systems
  • He Kong +8
  • Research Article

Cooperative Multiagent Reinforcement Learning Coupled With A* Search for Ship Multicabin Equipment Layout Considering Pipe Route

  • Apr 05, 2024
  • Journal of Ship Production and Design
  • Qiaoyu Zhang +1
  • Research Article
  • Citations2

Leveraging Organizational Hierarchy to Simplify Reward Design in Cooperative Multi-agent Reinforcement Learning

  • May 12, 2024
  • The International FLAIRS Conference Proceedings
  • Lixing Liu +2
  • Research Article

LLM Collaboration with Multi-Agent Reinforcement Learning

  • Mar 14, 2026
  • Shuo Liu +3
  • Research Article
  • Citations72

Reward shaping for knowledge-based multi-objective multi-agent reinforcement learning

  • Jan 01, 2018
  • The Knowledge Engineering Review
  • Patrick Mannion +3
  • Conference Article
  • Citations1

Centralized Critic per Knowledge for Cooperative Multi-Agent Game Environments

  • Oct 01, 2021
  • Thais Ferreira +2
  • Research Article
  • Citations2

Calculating the Reliability and Reputation of Agents in E-Commerce Multi-Agent Environments

  • Apr 28, 2014
  • Applied Mechanics and Materials
  • Elham Majd +5
  • PDF
  • Research Article
  • Citations2

Multiagent Cooperative Reinforcement Learning by Expert Agents (MCRLEA)

  • Jan 01, 2017
  • International Journal of Intelligent Information Systems
  • Deepak Annasaheb Vidhate
  • PDF
  • Research Article
  • Citations15

On Centralized Critics in Multi-Agent Reinforcement Learning

  • May 31, 2023
  • Journal of Artificial Intelligence Research
  • Xueguang Lyu +4
  • PDF
  • Research Article
  • Citations6

GHQ: grouped hybrid Q-learning for cooperative heterogeneous multi-agent reinforcement learning

  • Apr 23, 2024
  • Complex & Intelligent Systems
  • Xiaoyang Yu +4
  • Research Article
  • Citations2

A Coordination Optimization Framework for Multi-Agent Reinforcement Learning Based on Reward Redistribution and Experience Reutilization

  • Jun 09, 2025
  • Electronics
  • Bo Yang +7
  • Conference Article
  • Citations36

A Multiagent Reinforcement Learning approach for inverse kinematics of high dimensional manipulators with precision positioning

  • Jun 01, 2016
  • Yasmin Ansari +5
  • Conference Article

Cooperative multi-agent reinforcement learning based on online heuristic extraction

  • Jul 01, 2011
  • Jun Wu +3
  • Conference Article
  • Citations1

ADMN: Agent-Driven Modular Network for Dynamic Parameter Sharing in Cooperative Multi-Agent Reinforcement Learning

  • Aug 01, 2024
  • Yi Zhan +7
Cactus Communications logo

Copyright 2026 Cactus Communications. All rights reserved.