Home /Research /Multi-Robot Flocking Control Using Multi-Agent Twin Delayed Deep Deterministic Policy Gradient
SWARM

Multi-Robot Flocking Control Using Multi-Agent Twin Delayed Deep Deterministic Policy Gradient

Mario Salama Youssef, Nouran Adel Hassan, Ayman El-Badawy

Year
2022
Citations
5

Abstract

This paper proposes the use of Multi-Agent Twin Delayed Deep Deterministic Policy Gradient (MATD3) to resolve the issues accompanied with Multi-Agent Deep Deterministic Policy Gradient (MADDPG) such as the overestimation bias of a centralized critic when used to solve the multi-robot flocking control problem. A temporal difference error prioritized replay buffer is used alongside to achieve the most efficient learning process. In order to allow the robots to maintain a flock throughout a complex environment, an appropriate reward function was constructed, taking into consideration the following parameters: reaching the goal, maintaining a distance between the agents to ensure a stable connection while avoiding collision between one another, obstacle avoidance, and moving in a specific formation. Results of both algorithms are then compared to test the performance of the MATD3.

Keywords

Flocking (texture)Computer scienceObstacle avoidanceRobotObstacleReinforcement learningCollision avoidanceCollisionControl theory (sociology)Artificial intelligence

Related papers

Browse all SWARM papers