Deep Reinforcement Learning of Cooperative Control with Four Robotic Agents by MADDPG
Zhaoyang Wang, Renzhuo Wan, Xi Gui, Guopeng Zhou
- Year
- 2020
- Citations
- 8
Abstract
Due to the nature of complexity, inflexibility and non-robustness of classical cooperative control algorithms, the deep reinforcement learning has been widely researched and applied in collective and continuous behaviour control. Especially for multi-agents in real world, acquiring a full view world with a quick learning is still a great challenge. Inspired by Policy Gradient (PG) and its successors, a toy model with multi-agents by four two-dimensional manipulators environment is built based on physics engine-based MuJoCo. With a modified deep deterministic policy gradient algorithm and different credit strategies for individual agent, the cooperation and competition behaviour to target location between agents are studied. The experimental results show that each robot can complete the task with a negligible convergence effect, indicating that the MADDPG algorithm has a good performance in a complex environment, and successfully learn the strategy of multi-agent collaboration. However, with the instability of the environment caused by the increase in the number of agents, deep reinforcement learning has certain difficulties in the joint action space.
Keywords
Related papers
Statistical Learning Theory
Yuhai Wu, Vladimir Vapnik
1999
Artificial intelligence: a modern approach
1995
Applied Nonlinear Control
Jean-Jacques Slotine, Weiping Li
1991
A new optimizer using particle swarm theory
R.C. Eberhart, James Kennedy
2002