首页 /研究 /MCT-Grasp: A Novel Grasp Detection Using Multimodal Embedding and Convolutional Modulation Transformer
HRI

MCT-Grasp: A Novel Grasp Detection Using Multimodal Embedding and Convolutional Modulation Transformer

G.-H. Yang, Tong Jia, Yizhe Liu, Zhenghao Liu, Kaibo Zhang, Zhenjun Du

发表年份
2024
引用次数
7

摘要

For the foundation of human-robot intelligent interaction, it is crucial to endow robots with the ability to grasp accurately. Recently, some studies have introduced vision transformers (ViTs) into grasp detection networks to enhance detection accuracy. However, most of these works are initial attempts and do not specifically enhance learning spatial relationships and partial geometric information about objects. Therefore, this article proposes a novel transformer-based grasp network named MCT-Grasp, starting from the requirements of the grasping task. First, multimodal geometric positional embedding (MG-PE) is proposed to enhance the network’s ability to perceive spatial position. Then, a convolutional modulation (ConvMod) attention-based encoder is introduced, which can efficiently capture global and local input data features. Finally, the same-scale features from the encoder are integrated into the decoder to improve object edge localization and enrich the decoding information. The technique achieves an accuracy of 99.66% and 96.64% on the public datasets of Cornell and Jacquard, respectively. Moreover, the success rate reaches 91.67% in real-world grasping. Extensive experiments demonstrate that our method can achieve SOTA performance. Code and experimental video are available at <uri xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">https://github.com/YangXuanyi/MCT-Grasp</uri>.

关键词

GRASPEmbeddingComputer scienceTransformerArtificial intelligenceEngineeringElectrical engineeringVoltage

相关论文

查看 HRI 分类全部论文