Andrew Critch

University of California, Berkeley

Papers

2

Total Citations

11

H-Index

2

About

Andrew Critch is a leading researcher at the intersection of artificial intelligence safety, game theory, and human-AI alignment. His work critically examines how AI systems can infer and act upon human preferences, especially when those preferences are irrational or conflicting. In his highly cited 2021 paper, *"Human irrationality: both bad and good for reward inference,"* Critch explores how deviations from rational behavior—often dismissed as noise—can actually provide valuable signals for reward learning. This work has become foundational for researchers designing AI that robustly interprets human goals. Critch also made a landmark contribution to multi-agent alignment with his 2020 paper, *"Multi-Principal Assistance Games: Definition and Collegial Mechanisms."* Here, he introduced a novel framework for a single AI to assist multiple human principals with divergent values, cleverly circumventing classic impossibility results like Gibbard's theorem. This work has garnered attention for its practical approach to democratic AI governance. Beyond his publications, Critch is a co-founder of the AI safety organization *Encultured AI* and a former research scientist at DeepMind, where he helped shape the field's understanding of scalable oversight and value alignment.

Research Focus

Key Achievements

2
H-Index
2
Papers
11
Total Citations
6
Avg Citations/Paper
🏆 Most Cited Paper
Human irrationality: both bad and good for reward inference
7 citations · 2021
📈 Most Prolific Year: 2021 (1 Papers)
🤝 Key Collaborators: 6
🏛 Institutions: University of California, Berkeley

Top Papers

  1. 1
  2. 2

Key Collaborators

Contact & Links

Available for collaboration
Content generated · 12 days ago