Papers
11
Total Citations
658
H-Index
9
About
Qi Wu is a prominent AI researcher whose work sits at the intersection of computer vision, natural language processing, and embodied intelligence, with a particular focus on Vision-and-Language Navigation (VLN). His research addresses one of AI's most ambitious goals: building intelligent agents capable of understanding natural language instructions and navigating complex real-world environments autonomously. Wu's most influential contribution, the REVERIE benchmark (2020, 297 citations), challenged the field by introducing high-level, human-like instructions for remote embodied referring expression tasks in realistic indoor environments — pushing robots beyond simple room navigation toward meaningful object interaction. His foundational work on VLN (2018) helped establish the task's core framework, while his comprehensive survey (2022, 91 citations) has become an essential reference for researchers entering the field. His Object-and-Action Aware Model further refined how agents interpret visually-grounded instructions by distinguishing between object-centric and action-centric language cues. More recently, Wu has pioneered the integration of large vision-language models into navigation through NavGPT-2 (2024) and interactive prompting strategies, signaling his forward-thinking approach to leveraging generative AI for embodied reasoning. His body of work has significantly shaped how the research community conceptualizes human-robot collaboration through language.
Research Focus
Key Achievements
Top Papers
- 1REVERIE: Remote Embodied Visual Referring Expression in Real Indoor Environments297 citations · 2020
- 2Object-and-Action Aware Model for Visual Language Navigation98 citations · 2020
- 3
- 4
- 5
- 6
- 7
- 8
- 9
- 10Object-and-Action Aware Model for Visual Language Navigation7 citations · 2020