Vasil Khalidov
Papers
7
Total Citations
114
H-Index
5
About
Vasil Khalidov is a leading researcher at the intersection of human-robot interaction (HRI), multimodal perception, and self-supervised learning. His work fundamentally addresses how machines can understand and engage with humans in social, multi-party settings. Khalidov made seminal contributions to HRI by creating the **Vernissage Corpus** (2013, 33 citations), a benchmark dataset for conversational HRI that captures rich, real-behaving robot interactions with multiple humans. This work, alongside his studies on **engagement-based multi-party dialog** (2011, 28 citations) and **context-aware addressee estimation** (2013, 10 citations), established foundational methods for robots to detect visual focus of attention, recognize speakers, and decide when to engage users. His research on **finding audio-visual events** (2011, 22 citations) introduced novel multimodal clustering algorithms for detecting people who can be both seen and heard. More recently, Khalidov has pushed the frontier of AI with **V-JEPA 2** (2025, 3 citations), a self-supervised video model that learns to understand, predict, and plan by combining internet-scale video with minimal robot interaction data. This work represents a major step toward machines that can learn world models through observation, bridging perception and action. With over 100 citations across his key papers, Khalidov’s research continues to shape how robots perceive, interact with, and learn from the social world.
Research Focus
Key Achievements
Top Papers
- 1The vernissage corpus: A conversational Human-Robot-Interaction dataset33 citations · 2013
- 2Engagement-based Multi-party Dialog with a Humanoid Robot28 citations · 2011
- 3Finding audio-visual events in informal social gatherings22 citations · 2011
- 4THE VERNISSAGE CORPUS: A MULTIMODAL HUMAN-ROBOT-INTERACTION DATASET15 citations · 2012
- 5Context aware addressee estimation for human robot interaction10 citations · 2013
- 6
- 7