Dmitrii Krasheninnikov
Papers
1
Total Citations
15
H-Index
1
About
Dmitrii Krasheninnikov is a researcher whose work lies at the intersection of reinforcement learning (RL), AI safety, and value alignment. His most-cited paper, "Preferences Implicit in the State of the World" (2019, 15 citations), makes a foundational contribution by exposing a critical blind spot in reward specification: RL agents optimize only the features explicitly rewarded, remaining indifferent to everything else. Krasheninnikov argues that we must not only define what an agent should do but also the far larger space of what it should not do—a challenge he terms "implicit preferences." This insight has shaped how the AI safety community thinks about reward misspecification and the dangers of incomplete objective functions. His work highlights the subtle ways in which the state of the world encodes unstated preferences, urging researchers to consider the unintended consequences of narrow reward signals. Though early in his career, Krasheninnikov’s ideas have already influenced discussions on robust and aligned AI systems, making him a rising voice in the effort to ensure that advanced agents act in accordance with human values.
Research Focus
Key Achievements
Top Papers
- 1Preferences Implicit in the State of the World15 citations · 2019