PPO and SAC Reinforcement Learning Based Reference Compensation
Amir Hossein Dezhdar, D. Rahmati, Mohammad Mohtashami, Iman Sharifi
- 发表年份
- 2024
- 引用次数
- 2
摘要
In this paper, a reinforcement learning-based compensation method is used to control the UR5e robot, designed for applications such as rehabilitation. The robot is tasked with accurately reaching a fixed point in its workspace as well as performing smooth and precise tracking of predefined movement paths. Two reinforcement learning (RL) algorithms, Proximal Policy Optimization (PPO) and Soft Actor-Critic (SAC) are employed as the basis for the compensator. For evaluation, a simulation of the UR5e robot, a 6-DoF industrial robot, is used to follow reference paths, such as a square path. Based on the reference in the workspace, the required position for each joint is calculated using the inverse kinematics of the UR5e. Consequently, the reinforcement learning algorithm takes the error between each joint's reference and current state and computes a correction signal. This output is then used to adjust the reference point and generate an altered error, which is processed by the proportional-derivative (PD) controller to compute the velocity control signal provided to the UR5e.
关键词
相关论文
Statistical Learning Theory
Yuhai Wu, Vladimir Vapnik
1999
Artificial intelligence: a modern approach
1995
Applied Nonlinear Control
Jean-Jacques Slotine, Weiping Li
1991
A new optimizer using particle swarm theory
R.C. Eberhart, James Kennedy
2002