Papers
17
Total Citations
340
H-Index
6
About
Sean Kirmani is a robotics and AI researcher whose work sits at the intersection of vision-language models, spatial reasoning, and robotic learning. His most influential contribution, **SpatialVLM** (2024, 163 citations), addresses a critical gap in modern AI by endowing Vision-Language Models with 3D spatial reasoning capabilities — a foundational requirement for both visual question answering and real-world robotics deployment. Complementing this, his work on **Language to Rewards for Robotic Skill Synthesis** (2023) and **PromptBook for Manipulation Skills** (2024) demonstrates his sustained focus on leveraging large language models to bridge the gap between semantic reasoning and low-level robotic control. Kirmani has also contributed to large-scale robot deployment, including deep reinforcement learning systems for waste sorting across office building fleets and the **AutoRT** framework for orchestrating robotic agents at scale. His earlier work explored human-robot interaction through LED-based robot signaling and semantic mapping for navigation. Across his career, Kirmani has consistently tackled challenges of generalization, scalability, and real-world grounding — problems central to making robots genuinely useful. With over 300 cumulative citations, his research is shaping how the next generation of intelligent, language-guided robots will perceive and act in the physical world.
Research Focus
Key Achievements
Top Papers
- 1SpatialVLM: Endowing Vision-Language Models with Spatial Reasoning Capabilities163 citations · 2024
- 2
- 3Language to Rewards for Robotic Skill Synthesis38 citations · 2023
- 4
- 5
- 6
- 7PRISM: Pose Registration for Integrated Semantic Mapping5 citations · 2018
- 8
- 9
- 10PIVOT: Iterative Visual Prompting Elicits Actionable Knowledge for VLMs5 citations · 2024