Papers
7
Total Citations
245
H-Index
4
About
Yunzhi Zhang is a researcher whose work spans generative modeling, video prediction, and reinforcement learning, with notable contributions to the intersection of deep learning and sequential decision-making. Zhang is perhaps best known for co-developing **VideoGPT** (2021), a landmark architecture that applies VQ-VAE and transformer-based models to natural video generation using 3D convolutions and axial self-attention — a work that has garnered over 144 citations and helped establish scalable likelihood-based video generation as a viable research direction. Building on this foundation, Zhang contributed to **MaskViT** (2022), which demonstrated that masked visual pre-training of transformers yields strong video prediction models useful for embodied agent planning. Beyond generative modeling, Zhang has made meaningful contributions to reinforcement learning, including **Automatic Curriculum Learning through Value Disagreement** (2020), which addresses multi-goal RL by intelligently selecting training tasks, and work on asynchronous methods for model-based RL. Across these domains, Zhang's research reflects a consistent interest in making learning systems more efficient, scalable, and capable of anticipating future states — skills increasingly critical in robotics, planning, and AI-driven simulation.
Research Focus
Key Achievements
Top Papers
- 1VideoGPT: Video Generation using VQ-VAE and Transformers144 citations · 2021
- 2MaskViT: Masked Visual Pre-Training for Video Prediction45 citations · 2022
- 3Automatic Curriculum Learning through Value Disagreement34 citations · 2020
- 4Asynchronous Methods for Model-Based Reinforcement Learning13 citations · 2019
- 5
- 6
- 7VideoGen: Generative Modeling of Videos using VQ-VAE and Transformers2 citations · 2021