首页 /研究 /Multi-Robot Flocking Control Using Multi-Agent Twin Delayed Deep Deterministic Policy Gradient
SWARM

Multi-Robot Flocking Control Using Multi-Agent Twin Delayed Deep Deterministic Policy Gradient

Mario Salama Youssef, Nouran Adel Hassan, Ayman El-Badawy

发表年份
2022
引用次数
5

摘要

This paper proposes the use of Multi-Agent Twin Delayed Deep Deterministic Policy Gradient (MATD3) to resolve the issues accompanied with Multi-Agent Deep Deterministic Policy Gradient (MADDPG) such as the overestimation bias of a centralized critic when used to solve the multi-robot flocking control problem. A temporal difference error prioritized replay buffer is used alongside to achieve the most efficient learning process. In order to allow the robots to maintain a flock throughout a complex environment, an appropriate reward function was constructed, taking into consideration the following parameters: reaching the goal, maintaining a distance between the agents to ensure a stable connection while avoiding collision between one another, obstacle avoidance, and moving in a specific formation. Results of both algorithms are then compared to test the performance of the MATD3.

关键词

Flocking (texture)Computer scienceObstacle avoidanceRobotObstacleReinforcement learningCollision avoidanceCollisionControl theory (sociology)Artificial intelligence

相关论文

查看 SWARM 分类全部论文