Yanhao Zhang

Papers

1

Total Citations

2

H-Index

1

About

Yanhao Zhang is a rising researcher at the forefront of 3D perception, with a focus on leveraging multimodal large language models (MLLMs) to bridge the gap between 2D vision and 3D understanding. His seminal work, "LLMI3D: MLLM-based 3D Perception from a Single 2D Image" (2024), introduces a paradigm-shifting approach that enables robust 3D perception from just a single 2D image, addressing the critical limitations of traditional small models that struggle with generalization in open-world scenarios. This work is particularly impactful for applications in autonomous driving, augmented reality, robotics, and embodied intelligence, where accurate 3D understanding from minimal input is essential. Though early in its citation trajectory, the paper has already garnered attention for its innovative integration of MLLMs into 3D tasks. Zhang’s contributions are poised to redefine how machines perceive depth and spatial relationships, offering a scalable solution to one of computer vision’s most persistent challenges. His research not only advances theoretical understanding but also promises practical breakthroughs in real-world systems, marking him as a key figure in the next wave of intelligent perception.

Research Focus

Key Achievements

1
H-Index
1
Papers
2
Total Citations
2
Avg Citations/Paper
🏆 Most Cited Paper
LLMI3D: MLLM-based 3D Perception from a Single 2D Image
2 citations · 2024
📈 Most Prolific Year: 2024 (1 Papers)
🤝 Key Collaborators: 6

Top Papers

  1. 1

Key Collaborators

Contact & Links

Available for collaboration
Content generated · 10 days ago