Hongtao Wu

Tencent (China)

Papers

4

Total Citations

35

H-Index

4

About

Hongtao Wu is a researcher at the forefront of embodied AI, specializing in vision-language models, generative pre-training, and robot manipulation. His work addresses one of robotics' most pressing challenges: enabling robots to understand and interact with complex environments through multimodal learning. Wu's most influential contribution, "Vision-Language Foundation Models as Effective Robot Imitators" (19 citations), demonstrated how existing vision-language models can be efficiently fine-tuned to serve as capable robot imitators, lowering the barrier to deploying powerful AI in physical systems. Building on this foundation, his GR-2 framework pushed the boundaries further by pre-training a generalist robot agent on an unprecedented 38 million video clips and over 50 billion tokens sourced from the internet, achieving state-of-the-art versatility in robot manipulation tasks. His earlier work on large-scale video generative pre-training similarly established that visual representations learned from video data transfer effectively to robotic control. More recently, Wu has extended his expertise to legged locomotion, applying world model-based perception to help robots navigate challenging terrains. Collectively, his research demonstrates a consistent and impactful vision: harnessing internet-scale data and generative modeling to build more capable, generalizable robotic systems.

Research Focus

Key Achievements

4
H-Index
4
Papers
35
Total Citations
9
Avg Citations/Paper
🏆 Most Cited Paper
Vision-Language Foundation Models as Effective Robot Imitators
19 citations · 2023
📈 Most Prolific Year: 2023 (2 Papers)
🤝 Key Collaborators: 26
🏛 Institutions: Tencent (China)

Top Papers

  1. 1
  2. 2
  3. 3
  4. 4

Key Collaborators

Contact & Links

Available for collaboration
Content generated · 14 days ago