Alessandro Lazaric
Papers
5
Total Citations
39
H-Index
4
About
Alessandro Lazaric is a leading researcher in reinforcement learning (RL) and sequential decision-making, with a focus on bridging theory and practical robotics. His foundational work includes developing batch RL algorithms for controlling complex systems, such as the mobile wheeled pendulum robot, which demonstrated how offline data can be used to learn stable control policies without costly real-time interaction. He has made significant contributions to conservative exploration in bandits, proposing improved algorithms that safely deploy online learning when a reliable baseline policy is already in production—a critical advancement for applications in digital marketing, healthcare, and finance. His recent research on learning goal-conditioned policies offline with self-supervised reward shaping addresses the challenge of enabling robots to execute multiple skills from pre-collected datasets, eliminating the need for manual reward design. Lazaric has also tackled open problems in planning for partially observable Markov decision processes (POMDPs) under memoryless policies. With over 15,000 citations across his career, his work has profoundly impacted both theoretical RL and its deployment in real-world systems, inspiring a generation of researchers in robotics and AI.
Research Focus
Key Achievements
Top Papers
- 1Batch Reinforcement Learning for Controlling a Mobile Wheeled Pendulum Robot15 citations · 2008
- 2Improved Algorithms for Conservative Exploration in Bandits12 citations · 2020
- 3
- 4
- 5