Jasmine Hsu

Google (United States)

Papers

13

Total Citations

2,326

H-Index

11

About

Jasmine Hsu is a robotics and machine learning researcher whose work sits at the intersection of self-supervised learning, large-scale foundation models, and robotic control. She is perhaps best known for pioneering Time-Contrastive Networks, a self-supervised framework that learns rich visual representations directly from unlabeled multi-view video, enabling robots to imitate human object interactions without manually annotated data — work that has accumulated over 700 citations across its iterations. Her research trajectory then expanded toward grounding language in physical robot behavior, contributing to the landmark "Do As I Can, Not As I Say" paper (516 citations), which demonstrated how large language models can be paired with robotic affordances to execute complex, extended instructions. Hsu has also been a key contributor to Google's influential Robotics Transformer series — RT-1 and RT-2 — which showed that transformer-based models trained on large, diverse datasets can achieve unprecedented generalization in real-world robotic manipulation. Her involvement in the Open X-Embodiment initiative further reflects her commitment to community-scale data sharing across robot platforms. Collectively, her work has shaped the modern paradigm of large, generalizable robotic learning systems, earning well over 2,300 citations and establishing her as a significant voice in embodied AI research.

Research Focus

Key Achievements

11
H-Index
13
Papers
2,326
Total Citations
179
Avg Citations/Paper
🏆 Most Cited Paper
Time-Contrastive Networks: Self-Supervised Learning from Video
555 citations · 2018
📈 Most Prolific Year: 2023 (3 Papers)
🤝 Key Collaborators: 201
🏛 Institutions: Google (United States)

Top Papers

  1. 1
  2. 2
  3. 3
  4. 4
  5. 5
  6. 6
  7. 7
  8. 8
  9. 9
  10. 10

Key Collaborators

Contact & Links

Available for collaboration
Content generated · 13 days ago