6-DoF Pose Estimation Using an End-to-End Pipeline of Neural Network With BP-<i>n</i>-P
Mahfud Jiono, Hsien-I Lin, Wenhui Chen
- 发表年份
- 2024
- 引用次数
- 3
摘要
This article presents a single-shot, end-to-end deep learning architecture for object detection and 6-DoF pose estimation using an RGB image without postprocessing. Previous studies have focused on pose estimation using a two-stage pipeline consisting of a feature extraction network and a pose estimation stage. However, we propose a new network structure that consolidates these two stages into a single-stage architecture to enhance the performance. Our proposed framework involves an end-to-end pipeline that integrates a YOLO architecture with a backpropagating perspective-n-point (BP-n-P) layer to predict the object’s 6-DoF pose directly from an RGB image. The model aims to predict the 2-D image coordinates of the corners and center of the 3-D tight bounding box around the object in the image. Here, the network backpropagates the pose estimation error through the P-n-P solver to learn the correct object pose. The LineMod dataset, which includes 15,783 RGB images and 13 different object types, was used in the experiment. In the case of single-object-pose estimation, we obtained a mean pose accuracy of 71.27% and a mean pixel error of 9.16 pixels without any postprocessing or pose refinement. Our approach significantly outperformed other single-shot convolutional neural network (CNN) approaches without refinement conditions such as BB8, SSD-6D, YOLO-6D, PVNet, and GDR-Net. The proposed approach is suitable for robotic manipulation, self-driving cars, and augmented reality.
关键词
相关论文
Statistical Learning Theory
Yuhai Wu, Vladimir Vapnik
1999
Artificial intelligence: a modern approach
1995
Applied Nonlinear Control
Jean-Jacques Slotine, Weiping Li
1991
A new optimizer using particle swarm theory
R.C. Eberhart, James Kennedy
2002