Xiaohan Zhang
Papers
8
Total Citations
248
H-Index
6
About
Xiaohan Zhang is an emerging robotics and AI researcher whose work sits at the intersection of robot planning, embodied intelligence, and computer vision. Their most influential contributions focus on **task and motion planning (TAMP)** for service robots, with a particular emphasis on integrating large language models (LLMs) to enable commonsense reasoning in complex manipulation tasks — work that has already garnered 124 citations since 2023, reflecting rapid community uptake. Zhang has pioneered visually grounded approaches to TAMP, allowing robots to interpret and act upon visual scene understanding during long-horizon mobile manipulation tasks. Their contributions to **Embodied Question Answering** through the OpenEQA benchmark (46 citations) highlight a commitment to evaluating how foundation models can enable agents to genuinely understand and reason about physical environments. Beyond planning, Zhang has contributed to the field of **geo-localization**, authoring both original research and a comprehensive survey that maps the landscape of image and object localization techniques. Earlier work on 360° vision for telepresence robots further demonstrates a broad research vision spanning human-robot interaction. Collectively, Zhang's portfolio reflects a researcher shaping how intelligent robots perceive, reason, and act in unstructured real-world environments.
Research Focus
Key Achievements
Top Papers
- 1Task and Motion Planning with Large Language Models for Object Rearrangement124 citations · 2023
- 2OpenEQA: Embodied Question Answering in the Era of Foundation Models46 citations · 2024
- 3Image and Object Geo-Localization31 citations · 2023
- 4Visually Grounded Task and Motion Planning for Mobile Manipulation20 citations · 2022
- 5
- 6Visual and Object Geo-localization: A Comprehensive Survey8 citations · 2021
- 7Learning to Ground Objects for Robot Task and Motion Planning6 citations · 2022
- 8