首页 /研究 /Multi-Modal Attention-based Fusion Model for Semantic Segmentation of\n RGB-Depth Images
PERCEPTION

Multi-Modal Attention-based Fusion Model for Semantic Segmentation of\n RGB-Depth Images

Fahimeh Fooladgar, Shohreh Kasaei

发表年份
2019
引用次数
16
访问权限
开放获取

摘要

The 3D scene understanding is mainly considered as a crucial requirement in\ncomputer vision and robotics applications. One of the high-level tasks in 3D\nscene understanding is semantic segmentation of RGB-Depth images. With the\navailability of RGB-D cameras, it is desired to improve the accuracy of the\nscene understanding process by exploiting the depth features along with the\nappearance features. As depth images are independent of illumination, they can\nimprove the quality of semantic labeling alongside RGB images. Consideration of\nboth common and specific features of these two modalities improves the\nperformance of semantic segmentation. One of the main problems in RGB-Depth\nsemantic segmentation is how to fuse or combine these two modalities to achieve\nmore advantages of each modality while being computationally efficient.\nRecently, the methods that encounter deep convolutional neural networks have\nreached the state-of-the-art results by early, late, and middle fusion\nstrategies. In this paper, an efficient encoder-decoder model with the\nattention-based fusion block is proposed to integrate mutual influences between\nfeature maps of these two modalities. This block explicitly extracts the\ninterdependences among concatenated feature maps of these modalities to exploit\nmore powerful feature maps from RGB-Depth images. The extensive experimental\nresults on three main challenging datasets of NYU-V2, SUN RGB-D, and Stanford\n2D-3D-Semantic show that the proposed network outperforms the state-of-the-art\nmodels with respect to computational cost as well as model size. Experimental\nresults also illustrate the effectiveness of the proposed lightweight\nattention-based fusion model in terms of accuracy.\n

关键词

RGB color modelComputer scienceArtificial intelligenceSegmentationFeature (linguistics)Convolutional neural networkComputer visionEncoderModality (human–computer interaction)Pattern recognition (psychology)

相关论文

查看 PERCEPTION 分类全部论文