DTDE: A new cooperative multi-agent reinforcement learning framework
Guanghui Wen, Junjie Fu, Pengcheng Dai, Jialing Zhou
- Year
- 2021
- Citations
- 44
- Access
- Open access
Abstract
A significant body of work on reinforcement learning has been focused on the single-agent tasks where the agent aims to learn a policy that maximizes the cumulative reward in a dynamic environment.1Sutton R.S. Barto A.G. Reinforcement Learning: An Introduction. The MIT Press, 2018Google Scholar In the past decades, quite a few single-agent-based reinforcement learning algorithms have been developed in the literature.1Sutton R.S. Barto A.G. Reinforcement Learning: An Introduction. The MIT Press, 2018Google Scholar Yet, it is increasingly recognized that the single-agent-based reinforcement learning algorithms may fail to effectively handle large-scale optimization (decision) tasks with joint features. Within this context, cooperative multi-agent reinforcement learning (CMARL) algorithms have been proposed, where the agents aim to complete the multi-agent learning goal cooperatively through information exchange between neighboring agents. It has been witnessed in the past few years that CMARL algorithms have received increasing attention due to their broad applications in various fields, such as traffic signal control in intelligent transportation systems, energy management of smart grid, and coordination control of robot swarms. Compared with the single-agent reinforcement learning algorithm that considers only a single agent's state-action space, the joint state-action space of the CMARL algorithm grows exponentially as the number of agents increases.2Tan, M.. (1993). Multi-agent reinforcement learning: independent vs. cooperative agents. In Proc. 10th Int. Conf. Mach. Learn. 330-337.Google Scholar Therefore, CMARL encounters major challenges of algorithm complexity and scalability. Another challenge in designing an efficient CMARL algorithm is the partial observability of the environment in which each agent has to make its individual decisions based on the local observations. Within the context of CMARL, the independent Q-learning (IQL)-based algorithms, where each agent establishes a local Q-function by local state-action information, have been suggested and discussed in the literature.2Tan, M.. (1993). Multi-agent reinforcement learning: independent vs. cooperative agents. In Proc. 10th Int. Conf. Mach. Learn. 330-337.Google Scholar The advantages of several IQL-based CMARL algorithms compared with the independent MARL algorithms have been also examined.2Tan, M.. (1993). Multi-agent reinforcement learning: independent vs. cooperative agents. In Proc. 10th Int. Conf. Mach. Learn. 330-337.Google Scholar Although the IQL has good scalability, it sometimes cannot guarantee the collection of individual optimal actions of agents produced by local Q-functions equivalent to the optimal joint action, that is, the individual global max (IGM) principle may not be satisfied. Motivated partly by this observation, a new kind of CMARL paradigm based on the centralized training with decentralized execution (CTDE) mechanism has recently attracted significant attention. In CTDE, the agents' policies are trained with access to global information in a centralized manner and executed based only upon local observation in a decentralized way. Typical CTDE-based CMARL algorithms include value decomposition networks (VDN), among others.3Sunehag, P., Lever, G., Gruslys, A., et al. (2018). Value-decomposition networks for cooperative multi-agent learning based on team reward. In Proc. 17th Int. Conf. Auton. Agents MultiAgent Syst. 2085-2087.Google Scholar The aforementioned results advanced our knowledge of how to design CTDE-based algorithms for coping with CMARL problems. However, most of the above-mentioned CTDE-based algorithms are preliminarily focused on solving the CMARL problems in the absence of constraints on agents' actions, especially the inherent coupling (joint) constraints. However, due to the inherent complexity of large-scale CMARL problems, the feasible actions of an individual agent are generally affected by those of the other agen
Keywords
Related papers
Statistical Learning Theory
Yuhai Wu, Vladimir Vapnik
1999
Artificial intelligence: a modern approach
1995
Applied Nonlinear Control
Jean-Jacques Slotine, Weiping Li
1991
A new optimizer using particle swarm theory
R.C. Eberhart, James Kennedy
2002