Gengze Zhou
Papers
2
Total Citations
33
H-Index
2
About
Gengze Zhou is an emerging researcher at the forefront of embodied artificial intelligence, with a focused specialization in vision-and-language navigation (VLN) and large vision-language models (VLMs). His work addresses one of the most compelling challenges in modern AI: enabling autonomous agents to navigate unfamiliar environments by understanding and following natural language instructions. Zhou's research tackles the critical problem of generalization — specifically, how agents can transfer learned navigational skills to out-of-distribution scenes and bridge the gap between simulated training environments and real-world deployment. His most notable contribution, **NavGPT-2**, has garnered 31 citations since its 2024 publication, demonstrating significant early impact by unlocking sophisticated navigational reasoning capabilities within large vision-language models. His complementary work, **NaVid**, explores video-based VLM approaches to step-by-step navigation planning, further expanding the toolkit available for embodied AI research. Together, these contributions position Zhou as a promising voice in a rapidly evolving field where robotics, computer vision, and natural language processing converge. Researchers working on autonomous agents, human-robot interaction, or multimodal AI will find his work particularly relevant and forward-looking.
Research Focus
Key Achievements
Top Papers
- 1
- 2