MCT-Grasp: A Novel Grasp Detection Using Multimodal Embedding and Convolutional Modulation Transformer
G.-H. Yang, Tong Jia, Yizhe Liu, Zhenghao Liu, Kaibo Zhang, Zhenjun Du
- 发表年份
- 2024
- 引用次数
- 7
摘要
For the foundation of human-robot intelligent interaction, it is crucial to endow robots with the ability to grasp accurately. Recently, some studies have introduced vision transformers (ViTs) into grasp detection networks to enhance detection accuracy. However, most of these works are initial attempts and do not specifically enhance learning spatial relationships and partial geometric information about objects. Therefore, this article proposes a novel transformer-based grasp network named MCT-Grasp, starting from the requirements of the grasping task. First, multimodal geometric positional embedding (MG-PE) is proposed to enhance the network’s ability to perceive spatial position. Then, a convolutional modulation (ConvMod) attention-based encoder is introduced, which can efficiently capture global and local input data features. Finally, the same-scale features from the encoder are integrated into the decoder to improve object edge localization and enrich the decoding information. The technique achieves an accuracy of 99.66% and 96.64% on the public datasets of Cornell and Jacquard, respectively. Moreover, the success rate reaches 91.67% in real-world grasping. Extensive experiments demonstrate that our method can achieve SOTA performance. Code and experimental video are available at <uri xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">https://github.com/YangXuanyi/MCT-Grasp</uri>.
关键词
相关论文
Statistical Learning Theory
Yuhai Wu, Vladimir Vapnik
1999
Artificial intelligence: a modern approach
1995
Applied Nonlinear Control
Jean-Jacques Slotine, Weiping Li
1991
A new optimizer using particle swarm theory
R.C. Eberhart, James Kennedy
2002