About

Yunzhi Zhang is a researcher whose work spans generative modeling, video prediction, and reinforcement learning, with notable contributions to the intersection of deep learning and sequential decision-making. Zhang is perhaps best known for co-developing **VideoGPT** (2021), a landmark architecture that applies VQ-VAE and transformer-based models to natural video generation using 3D convolutions and axial self-attention — a work that has garnered over 144 citations and helped establish scalable likelihood-based video generation as a viable research direction. Building on this foundation, Zhang contributed to **MaskViT** (2022), which demonstrated that masked visual pre-training of transformers yields strong video prediction models useful for embodied agent planning. Beyond generative modeling, Zhang has made meaningful contributions to reinforcement learning, including **Automatic Curriculum Learning through Value Disagreement** (2020), which addresses multi-goal RL by intelligently selecting training tasks, and work on asynchronous methods for model-based RL. Across these domains, Zhang's research reflects a consistent interest in making learning systems more efficient, scalable, and capable of anticipating future states — skills increasingly critical in robotics, planning, and AI-driven simulation.

Research Focus

Key Achievements

4
H-Index
7
Papers
245
Total Citations
35
Avg Citations/Paper
🏆 Most Cited Paper
VideoGPT: Video Generation using VQ-VAE and Transformers
144 citations · 2021
📈 Most Prolific Year: 2021 (2 Papers)
🤝 Key Collaborators: 18
🏛 Institutions: University of California, Berkeley, Nanjing University of Aeronautics and Astronautics, Commercial Aircraft Corporation of China (China)

Top Papers

  1. 1
  2. 2
  3. 3
  4. 4
  5. 5
  6. 6
  7. 7

Key Collaborators

Contact & Links

Available for collaboration
Content generated · 13 days ago