Papers

12

Total Citations

6,809

H-Index

12

About

John Schulman is a prominent artificial intelligence researcher whose work spans reinforcement learning, robotic motion planning, and deep learning. He is perhaps best known for developing Trust Region Policy Optimization (TRPO), a landmark 2015 algorithm that introduced theoretically grounded, monotonically improving policy updates — a breakthrough that reshaped modern reinforcement learning and has accumulated over 3,100 citations. Complementing this, his Generalized Advantage Estimation (GAE) paper offered elegant solutions to the variance-bias tradeoff in policy gradient methods, earning nearly 1,750 citations and becoming a staple technique in the field. Earlier in his career, Schulman made significant contributions to robotic motion planning, developing sequential convex optimization approaches for efficient, collision-free trajectory generation — work that attracted hundreds of citations and laid important groundwork for practical robot control. He has also explored dexterous manipulation, learning from demonstrations, and hierarchical meta-learning, reflecting a remarkably broad research vision. His contributions to deep reinforcement learning, particularly in making policy optimization both principled and scalable, have profoundly influenced the training of modern AI systems — including large language models — cementing his reputation as one of the field's most impactful researchers.

Research Focus

Key Achievements

12
H-Index
12
Papers
6,809
Total Citations
567
Avg Citations/Paper
🏆 Most Cited Paper
Trust Region Policy Optimization
3,141 citations · 2015
📈 Most Prolific Year: 2015 (3 Papers)
🤝 Key Collaborators: 26
🏛 Institutions: University of California, Berkeley, OpenAI (United States)

Top Papers

  1. 1
    Trust Region Policy Optimization
    3,141 citations · 2015
  2. 2
  3. 3
  4. 4
  5. 5
  6. 6
  7. 7
  8. 8
  9. 9
  10. 10

Key Collaborators

Contact & Links

Available for collaboration
Content generated · 13 days ago