EffoNAV: An Effective Foundation-Model-Based Visual Navigation Approach in Challenging Environment
Wangtian Shen, Pengfei Gu, Haijian Qin, Ziyang Meng
- 发表年份
- 2025
- 引用次数
- 1
摘要
Image-goal navigation is a critical task in autonomous visual navigation, requiring the robot to navigate to a target localization specified by an image. Previous works using data-driven methods achieve great success while they mostly leverage simple network architecture and train it from scratch, which limits the navigation performance in challenging situations, such as multiple turns or varying lighting conditions. In this paper, we thoroughly analyze the essential features for visual navigation and design an effective network to achieve optimal navigation performance. In particular, we leverage a pretrained foundation model for feature extraction, introduce cross attention for goal encoding and propose a token attention mechanism to dynamically assign weights to different tokens. The proposed model achieves excellent navigation performance in unseen environments. Experiments in real world demonstrate that our method achieves a success rate of 87%, 40% improvement over the state-of-the-art methods. For more information about the code, see <uri xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">https://github.com/robotnav-bot/EffoNAV</uri>.
关键词
相关论文
Statistical Learning Theory
Yuhai Wu, Vladimir Vapnik
1999
Artificial intelligence: a modern approach
1995
Fractional Differential Equations
Igor Podlubný
2025
Applied Nonlinear Control
Jean-Jacques Slotine, Weiping Li
1991