John Schulman
Papers
12
Total Citations
6,809
H-Index
12
About
John Schulman is a prominent artificial intelligence researcher whose work spans reinforcement learning, robotic motion planning, and deep learning. He is perhaps best known for developing Trust Region Policy Optimization (TRPO), a landmark 2015 algorithm that introduced theoretically grounded, monotonically improving policy updates — a breakthrough that reshaped modern reinforcement learning and has accumulated over 3,100 citations. Complementing this, his Generalized Advantage Estimation (GAE) paper offered elegant solutions to the variance-bias tradeoff in policy gradient methods, earning nearly 1,750 citations and becoming a staple technique in the field. Earlier in his career, Schulman made significant contributions to robotic motion planning, developing sequential convex optimization approaches for efficient, collision-free trajectory generation — work that attracted hundreds of citations and laid important groundwork for practical robot control. He has also explored dexterous manipulation, learning from demonstrations, and hierarchical meta-learning, reflecting a remarkably broad research vision. His contributions to deep reinforcement learning, particularly in making policy optimization both principled and scalable, have profoundly influenced the training of modern AI systems — including large language models — cementing his reputation as one of the field's most impactful researchers.
Research Focus
Key Achievements
Top Papers
- 1Trust Region Policy Optimization3,141 citations · 2015
- 2High-Dimensional Continuous Control Using Generalized Advantage Estimation1,750 citations · 2015
- 3
- 4
- 5Learning from Demonstrations Through the Use of Non-rigid Registration139 citations · 2016
- 6
- 7Meta Learning Shared Hierarchies117 citations · 2017
- 8
- 9
- 10