Agrim Gupta
Papers
7
Total Citations
438
H-Index
7
About
Agrim Gupta is a researcher at the intersection of machine learning, robotics, and computer vision, with a focus on developing intelligent systems capable of navigating and manipulating complex real-world environments. His early and highly influential work, "Social GAN" (2018, 180 citations), introduced generative adversarial networks to model the multimodal nature of human motion trajectories — a breakthrough for autonomous vehicles and social robots operating in human-centric spaces. Building on this foundation, Gupta has made substantial contributions to generalizable robot learning, co-authoring "Open X-Embodiment" (2024, 119 citations), a landmark effort to consolidate diverse robotic datasets and train large-scale, transferable models analogous to foundation models in NLP and vision. His work on "VIMA" (2022, 65 citations) pioneered multimodal prompt-based robot manipulation, while "MaskViT" (2022, 45 citations) advanced video prediction through masked visual pre-training with transformers. Projects like "MetaMorph" and "RoboCat" further demonstrate his commitment to universal, self-improving robotic controllers. Across his career, Gupta's research consistently pushes toward scalable, generalizable intelligence — bridging perception, prediction, and physical interaction in embodied AI systems.
Research Focus
Key Achievements
Top Papers
- 1
- 2
- 3VIMA: General Robot Manipulation with Multimodal Prompts65 citations · 2022
- 4MaskViT: Masked Visual Pre-Training for Video Prediction45 citations · 2022
- 5MetaMorph: Learning Universal Controllers with Transformers12 citations · 2022
- 6RoboCat: A Self-Improving Generalist Agent for Robotic Manipulation9 citations · 2023
- 7