Papers
3
Total Citations
10
H-Index
2
About
Arunesh Sinha is a researcher whose work lies at the critical intersection of reinforcement learning (RL) and AI safety, with a particular focus on security vulnerabilities and constrained decision-making. His most prominent contributions expose a dangerous new attack vector in offline RL, where agents learn from pre-collected datasets rather than live interaction. Through his work on the BAFFLE framework, Sinha demonstrates how backdoors can be stealthily hidden within offline RL datasets, allowing an attacker to trigger malicious behavior in trained agents while maintaining normal performance otherwise. This research, accumulating 8 citations across its iterations, highlights a fundamental security gap in the increasingly popular offline RL paradigm. Sinha also addresses the challenge of safety in long-horizon tasks through constrained hierarchical RL, proposing methods to handle richly constrained problems that extend beyond simple short-horizon scenarios. His work is particularly relevant as RL systems move from simulated environments to real-world applications where both security and safety are paramount. By revealing how data providers could weaponize shared datasets, Sinha’s research serves as a crucial warning and foundation for developing more robust, trustworthy RL systems.
Research Focus
Key Achievements
Top Papers
- 1Baffle: Hiding Backdoors in Offline Reinforcement Learning Datasets5 citations · 2024
- 2BAFFLE: Hiding Backdoors in Offline Reinforcement Learning Datasets3 citations · 2022
- 3