首页 /研究 /Transformer-Based Model for Monocular Visual Odometry: A Video Understanding Approach

PERCEPTION

Transformer-Based Model for Monocular Visual Odometry: A Video Understanding Approach

André O. Françani, Marcos R. O. A. Maximo

发表年份: 2023
访问权限: 开放获取

摘要

Estimating the camera's pose given images from a single camera is a traditional task in mobile robots and autonomous vehicles. This problem is called monocular visual odometry and often relies on geometric approaches that require considerable engineering effort for a specific scenario. Deep learning methods have been shown to be generalizable after proper training and with a large amount of available data. Transformer-based architectures have dominated the state-of-the-art in natural language processing and computer vision tasks, such as image and video understanding. In this work, we deal with the monocular visual odometry as a video understanding task to estimate the 6 degrees of freedom of a camera's pose. We contribute by presenting the TSformer-VO model based on spatio-temporal self-attention mechanisms to extract features from clips and estimate the motions in an end-to-end manner. Our approach achieved competitive state-of-the-art performance compared with geometry-based and deep learning-based methods on the KITTI visual odometry dataset, outperforming the DeepVO implementation highly accepted in the visual odometry community. The code is publicly available at https://github.com/aofrancani/TSformer-VO.

关键词

cs.CVcs.AIcs.RO

Transformer-Based Model for Monocular Visual Odometry: A Video Understanding Approach

摘要

关键词

相关论文

如何缓解越野环境中语义分割的分布偏移

基于原型模糊推理与证据融合的不确定性引导工业机器人可进化识别框架

基于点云配准的非破坏性高分辨率涂层厚度三维扫描测量

迈向智能机器人时代：用于高级感知系统的多模态柔性触觉传感器