Reinforcement learning, a machine learning technique, has already shown advancement in solving complex sequential decision-making problems. Many tasks involve multiple agents and require sequential decision-making policies to achieve common goals, such as warehouse automation, autonomous driving, and game-playing. To derive the policies for all agents, these problems can be modelled as multi-agent systems and addressed by multi-agent reinforcement learning (MARL). However, optimizing policies in multi-agent scenarios presents significant challenges due to the intricate behaviours of multiple agents and the non-stationary nature of the environments’ complex dynamics. Firstly, optimizing policies in multi-agent scenarios presents significant challenges due to the complexity of multi-agent behaviours, especially in partially observable environments. Also, the dynamic nature of agents’ behaviours and their interactions with other agents lead to changes in the environment’s states and agents’ observations over time, which is more complicated in open environments. Furthermore, the need to balance individual and collective objectives in certain real-world multi-agent environments also complicates the decision-making process. This doctoral thesis aims to tackle the following three fundamental multi-agent research problems and propose a solution for each. Our research covers from theoretical analysis to practical application. We start by studying the problem of learning an efficient policy under partially observable environments. We notice that a group of agents cooperates to com- pete against another group of agents (named opponents), with the limitation that information about the opponents is inaccessible due to partial observability. To address this issue, we propose a novel multi-agent distributional actor-critic algorithm to achieve speculative opponent modeling with purely local information (i.e., the controlled agent’s observations, actions, and rewards). The actor maintains a speculated belief of the opponents, which we call the speculative opponent models, to predict opponents’ actions using local observations and make decisions accordingly. The distributional critic models the return distribution of the policy. It reflects the quality of the actor and thus can guide the training of the imaginary opponent model that the actor relies on. Extensive experiments confirm that our method successfully models opponents’ behaviours without their data and delivers superior performance against baseline methods with a faster convergence speed. Furthermore, in some environments, teammates’ numbers and policies change with the market, requiring workers to adapt to perform different sets of tasks across time. To solve this issue, we propose an RL-based method that allows the controlled agent to collaborate with dynamic teammates in open environments. The controlled agent maintains a dual teamwork situation inference model to capture the current teamwork state and facilitate reasonable decision-making under partial observability. Considering the dynamic types of teammates, we first leverage the Chinese Restaurant Process-based model to categorize versatile teammate policies into distinct clusters, which improves the efficiency of identifying current teamwork situations. Next, to model the heterogeneous relationships among agents and accommodate the varying number of teammates, we apply heterogeneous graph attention neural networks to learn the representation of the teamwork situation. Extensive experiments confirm that our method outperforms state-of-the-art baselines with faster convergence in various ad hoc teamwork tasks. Lastly, in certain real-world applications, such as routing problems and warehouse management, decision-makers must balance overall benefits with individual fairness among agents. Achieving both learning efficiency and fairness concurrently presents a complex, multi-objective, joint-policy optimization challenge. Moreover, existing approaches are predominantly confined to simulation environments. To address the above issues, we present a pioneering MARL approach to balance individual and collective objectives for agents collaboration. Extensive experiments on synthetic and real-world datasets demonstrate that our method not only surpasses the state-of-the-art DRL methods, but also exhibits advantages over established heuristics, notably in achieving significantly faster optimization speeds. This method underlines the essential integration of fairness into real-world applications, representing a major advancement in creating equitable and efficient logistics solutions. To conclude, this doctoral thesis investigates three fundamental multi-agent decision-making research problems that are ubiquitous and unsolved. The proposed three MARL methodology solutions achieve efficient policy training and performance for agents in multi-agent environments with uncertainties raised by the partially observable, the open environment, and the individual-collective objectives of MARL. This thesis delves into novel designs of various critical components of MARL, including MDP formulations, policy networks, training algorithms, and inference methods. These contributions significantly elevate the effectiveness and efficiency of cooperative MARL, establishing new performance benchmarks.
Read more