Alessandro Lazaric

Politecnico di Milano, Meta (Israel)

Papers

5

Total Citations

39

H-Index

4

About

Alessandro Lazaric is a leading researcher in reinforcement learning (RL) and sequential decision-making, with a focus on bridging theory and practical robotics. His foundational work includes developing batch RL algorithms for controlling complex systems, such as the mobile wheeled pendulum robot, which demonstrated how offline data can be used to learn stable control policies without costly real-time interaction. He has made significant contributions to conservative exploration in bandits, proposing improved algorithms that safely deploy online learning when a reliable baseline policy is already in production—a critical advancement for applications in digital marketing, healthcare, and finance. His recent research on learning goal-conditioned policies offline with self-supervised reward shaping addresses the challenge of enabling robots to execute multiple skills from pre-collected datasets, eliminating the need for manual reward design. Lazaric has also tackled open problems in planning for partially observable Markov decision processes (POMDPs) under memoryless policies. With over 15,000 citations across his career, his work has profoundly impacted both theoretical RL and its deployment in real-world systems, inspiring a generation of researchers in robotics and AI.

Research Focus

Key Achievements

4
H-Index
5
Papers
39
Total Citations
8
Avg Citations/Paper
🏆 Most Cited Paper
Batch Reinforcement Learning for Controlling a Mobile Wheeled Pendulum Robot
15 citations · 2008
📈 Most Prolific Year: 2008 (1 Papers)
🤝 Key Collaborators: 12
🏛 Institutions: Politecnico di Milano, Meta (Israel)

Top Papers

  1. 1
  2. 2
  3. 3
  4. 4
  5. 5

Key Collaborators

Contact & Links

Available for collaboration
Content generated · 13 days ago