Bilateral Cross-Modal Fusion Network for Robot Grasp Detection
Qiang Zhang, Xueying Sun, Mingmin Liu, Zhiwei Fan
- 发表年份
- 2023
- 引用次数
- 2
- 访问权限
- 开放获取
摘要
In the field of vision-based robot grasping, effectively leveraging RGB and depth information to accurately determine the position and pose of a target is a critical issue. To address this challenge, we propose a tri-stream cross-modal fusion architecture for 2-DoF visual grasp detection. This architecture facilitates the interaction of RGB and depth bilateral information and is designed to efficiently aggregate multiscale information. Our novel modal interaction module (MIM) with spatial-wise cross-attention algorithm adaptively captures cross-modal feature information. Meanwhile, the channel interaction modules (CIM) further enhance the aggregation of different modal streams. In addition, we efficiently aggregate global multiscale information through a hierarchical structure with skipping connections. To evaluate the performance of our proposed method, we conduct validation experiments on standard public datasets and real robot grasping experiments. We achieve the image-wise detection accuracy of 99.4% and 96.7% on Cornell and Jacquard datasets respectively. The object-wise detection accuracy reaches 97.8% and 94.6% on the same datasets. Furthermore, physical experiments using the 6-DoF Elite robot demonstrate a success rate of 94.5%. These experiments highlight the superior accuracy of our proposed method.
关键词
相关论文
Statistical Learning Theory
Yuhai Wu, Vladimir Vapnik
1999
Artificial intelligence: a modern approach
1995
Applied Nonlinear Control
Jean-Jacques Slotine, Weiping Li
1991
A new optimizer using particle swarm theory
R.C. Eberhart, James Kennedy
2002