Stefan Welker
Papers
7
Total Citations
612
H-Index
5
About
Stefan Welker is a leading roboticist whose work sits at the intersection of foundation models and embodied AI, pioneering how large-scale pretrained models can transfer knowledge from the internet into the physical world. His most impactful contribution is **RT-2** (267 citations), a vision-language-action model that directly incorporates web-scale training data into end-to-end robotic control, enabling robots to generalize to novel objects and environments and even perform emergent semantic reasoning. This breakthrough builds on his earlier **Socratic Models** framework (171 citations), which composes zero-shot multimodal reasoning by bridging the gap between vision-language and large language models. Welker also introduced **Transporter Networks** (100 citations), a novel architecture that rearranges deep features to infer spatial displacements for precise manipulation, and demonstrated how **self-supervised deep reinforcement learning** can learn synergies between pushing and grasping (53 citations). His recent work on **AutoRT** and **Gemini Robotics** (2024-2025) pushes toward large-scale orchestration of embodied agents, leveraging foundation models to reason about tasks and scale data collection in the real world. Through these contributions, Welker has become a key figure in making robots that can reason, adapt, and act intelligently in unstructured environments.
Research Focus
Key Achievements
Top Papers
- 1RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control267 citations · 2023
- 2Socratic Models: Composing Zero-Shot Multimodal Reasoning with Language171 citations · 2022
- 3Transporter Networks: Rearranging the Visual World for Robotic\n Manipulation100 citations · 2020
- 4
- 5
- 6Gemini Robotics: Bringing AI into the Physical World4 citations · 2025
- 7