Home /Research /Enhancing position-based visual servoing performance through transformer-based acceleration-level reinforcement learning
MANIPULATION

Enhancing position-based visual servoing performance through transformer-based acceleration-level reinforcement learning

Wenkai Chen, Siqin Wang, Pengfei Wang, Ye Yuan, Tao Wu, Jianwei Zhang, Qingdu Li

Year
2025
Citations
2
Access
Open access

Abstract

Abstract Visual servoing is a fundamental approach for robotic manipulation that relies on visual feedback to precisely control robot motion. Most methods are capable of generating velocity control signals to guide the camera to the desired position and orientation, which often exhibit limitations in dynamic responsiveness and robustness against noise and unmodeled dynamics. This paper presents an innovative acceleration-level position-based visual servoing control framework enhanced by deep reinforcement learning (DRL) integrated with Transformer-based temporal sequence processing. The essence of the method comprises two key elements: First, the controller retains the theoretical approach of position-based visual servoing in its design, ensuring transparency and a guaranteed performance baseline. Second, considering the temporal characteristics of servoing control, a Transformer-based actor-critic architecture within a Proximal Policy Optimization (PPO) reinforcement learning scheme is proposed to improve the learning efficiency and performance. Comprehensive experiments are conducted in both simulation and real robot scenarios. The results reveal that, compared with traditional velocity-level controllers, the proposed method demonstrates superior dynamic characteristics, enhanced tracking performance, and diminished sensitivity to noise in Cartesian space.

Keywords

Visual servoingReinforcement learningReinforcementTransformerPosition (finance)Computer scienceComputational intelligenceArtificial intelligenceAccelerationComputer vision

Related papers

Browse all MANIPULATION papers