首页 /研究 /Bilateral Cross-Modal Fusion Network for Robot Grasp Detection
MANIPULATION

Bilateral Cross-Modal Fusion Network for Robot Grasp Detection

Qiang Zhang, Xueying Sun, Mingmin Liu, Zhiwei Fan

发表年份
2023
引用次数
2
访问权限
开放获取

摘要

In the field of vision-based robot grasping, effectively leveraging RGB and depth information to accurately determine the position and pose of a target is a critical issue. To address this challenge, we propose a tri-stream cross-modal fusion architecture for 2-DoF visual grasp detection. This architecture facilitates the interaction of RGB and depth bilateral information and is designed to efficiently aggregate multiscale information. Our novel modal interaction module (MIM) with spatial-wise cross-attention algorithm adaptively captures cross-modal feature information. Meanwhile, the channel interaction modules (CIM) further enhance the aggregation of different modal streams. In addition, we efficiently aggregate global multiscale information through a hierarchical structure with skipping connections. To evaluate the performance of our proposed method, we conduct validation experiments on standard public datasets and real robot grasping experiments. We achieve the image-wise detection accuracy of 99.4% and 96.7% on Cornell and Jacquard datasets respectively. The object-wise detection accuracy reaches 97.8% and 94.6% on the same datasets. Furthermore, physical experiments using the 6-DoF Elite robot demonstrate a success rate of 94.5%. These experiments highlight the superior accuracy of our proposed method.

关键词

Computer scienceArtificial intelligenceRobotGRASPRGB color modelAggregate (composite)ModalComputer visionFeature (linguistics)Pattern recognition (psychology)

相关论文

查看 MANIPULATION 分类全部论文