Papers
2
Total Citations
21
H-Index
2
About
Peng Gao is an emerging researcher at the intersection of robotics, computer vision, and large language models, with a focus on enabling more capable and generalizable robotic manipulation systems. His work addresses a critical gap in modern robotics: while multi-modal large language models (MLLMs) have transformed how robots interpret natural language, they often lack the domain-specific knowledge required for precise physical interaction. Gao's most-cited contribution, **ManipVQA** (2024, 18 citations), tackles this challenge directly by injecting robotic affordance and physically grounded information into MLLMs, meaningfully advancing robots' capacity to reason about and execute manipulation tasks. Building on this foundation, his more recent work **UniAff** (2025) proposes a unified representation framework that integrates 3D object-centric manipulation with affordance reasoning and task understanding, leveraging vision-language models to better capture underlying motion constraints. Together, these contributions reflect Gao's broader research vision: bridging the gap between high-level language understanding and low-level physical grounding in robotic systems. Though early in his career, his growing citation record signals increasing recognition from the robotics and AI research communities.
Research Focus
Key Achievements
Top Papers
- 1
- 2