Home /Research /Target Grasping and Multi-modal Interaction System Based on Pepper Robot
MANIPULATION

Target Grasping and Multi-modal Interaction System Based on Pepper Robot

Kuozhan Wang, Aihui Wang, Yan Wang, Xuebing Yue, Junjie Xie, Yangyang Wang

Year
2024
Citations
1

Abstract

Pepper robot integrates an RGB-D camera, voice system, and control system, and has been widely used in homes, shopping malls, and hotels. However, existing Pepper robots are mostly based on preset programs and have limited flexibility. To improve the intelligence level of robots, this paper presents a multi-modal interactive system that integrates YOLOv8 and the large language model (LLM). Firstly, to allow the robot to recognize and grasp target objects in the home environment, YOLOv8 target detection is integrated into the robot vision system based on the robot operating system (ROS), and the coordinates of the object are obtained by combining the depth camera positioning algorithm. Aiming at the limitation of the detection distance of the structured light depth camera, an odometer coordinate compensation strategy is designed to ensure that the object can stably obtain its coordinates when it is in the working space of the robot arm. To further enhance the intelligent interaction ability of the robot, LLM is integrated into the robot system, so that the robot can program independently to interact with the user according to the user’s intention. Finally, the feasibility of the proposed method is verified by experiments and the grasping and intelligent interaction of the target object are realized.

Keywords

ModalComputer scienceRobotHuman–computer interactionPepperArtificial intelligenceComputer security

Related papers

Browse all MANIPULATION papers