首页 /研究 /Simultaneous Multiple Object Detection and Pose Estimation using 3D Model Infusion with Monocular Vision

PERCEPTION

Simultaneous Multiple Object Detection and Pose Estimation using 3D Model Infusion with Monocular Vision

Congliang Li, Shijie Sun, Xiangyu Song, Huansheng Song, Naveed Akhtar, Ajmal Saeed Mian

发表年份: 2022
访问权限: 开放获取

摘要

Multiple object detection and pose estimation are vital computer vision tasks. The latter relates to the former as a downstream problem in applications such as robotics and autonomous driving. However, due to the high complexity of both tasks, existing methods generally treat them independently, which is sub-optimal. We propose simultaneous neural modeling of both using monocular vision and 3D model infusion. Our Simultaneous Multiple Object detection and Pose Estimation network (SMOPE-Net) is an end-to-end trainable multitasking network with a composite loss that also provides the advantages of anchor-free detections for efficient downstream pose estimation. To enable the annotation of training data for our learning objective, we develop a Twin-Space object labeling method and demonstrate its correctness analytically and empirically. Using the labeling method, we provide the KITTI-6DoF dataset with $\sim7.5$K annotated frames. Extensive experiments on KITTI-6DoF and the popular LineMod datasets show a consistent performance gain with SMOPE-Net over existing pose estimation methods. Here are links to our proposed SMOPE-Net, KITTI-6DoF dataset, and LabelImg3D labeling tool.

关键词

cs.CV

Simultaneous Multiple Object Detection and Pose Estimation using 3D Model Infusion with Monocular Vision

摘要

关键词

相关论文

如何缓解越野环境中语义分割的分布偏移

基于原型模糊推理与证据融合的不确定性引导工业机器人可进化识别框架

基于点云配准的非破坏性高分辨率涂层厚度三维扫描测量

迈向智能机器人时代：用于高级感知系统的多模态柔性触觉传感器