MCT-Grasp: A Novel Grasp Detection Using Multimodal Embedding and Convolutional Modulation Transformer
G.-H. Yang, Tong Jia, Yizhe Liu, Zhenghao Liu, Kaibo Zhang, Zhenjun Du
- Year
- 2024
- Citations
- 7
Abstract
For the foundation of human-robot intelligent interaction, it is crucial to endow robots with the ability to grasp accurately. Recently, some studies have introduced vision transformers (ViTs) into grasp detection networks to enhance detection accuracy. However, most of these works are initial attempts and do not specifically enhance learning spatial relationships and partial geometric information about objects. Therefore, this article proposes a novel transformer-based grasp network named MCT-Grasp, starting from the requirements of the grasping task. First, multimodal geometric positional embedding (MG-PE) is proposed to enhance the network’s ability to perceive spatial position. Then, a convolutional modulation (ConvMod) attention-based encoder is introduced, which can efficiently capture global and local input data features. Finally, the same-scale features from the encoder are integrated into the decoder to improve object edge localization and enrich the decoding information. The technique achieves an accuracy of 99.66% and 96.64% on the public datasets of Cornell and Jacquard, respectively. Moreover, the success rate reaches 91.67% in real-world grasping. Extensive experiments demonstrate that our method can achieve SOTA performance. Code and experimental video are available at <uri xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">https://github.com/YangXuanyi/MCT-Grasp</uri>.
Keywords
Related papers
Statistical Learning Theory
Yuhai Wu, Vladimir Vapnik
1999
Artificial intelligence: a modern approach
1995
Applied Nonlinear Control
Jean-Jacques Slotine, Weiping Li
1991
A new optimizer using particle swarm theory
R.C. Eberhart, James Kennedy
2002