Any-point Trajectory Modeling for Policy Learning
Xingyu Lin, John So, Kai Chen, Qi Dou, Yang Gao, Pieter Abbeel
- Year
- 2024
- Citations
- 40
- Access
- Open access
Abstract
Learning from demonstration is a powerful method for teaching robots new skills, and having more demonstration data often improves policy learning.However, the high cost of collecting demonstration data is a significant bottleneck.Videos, as a rich data source, contain knowledge of behaviors, physics, and semantics, but extracting control-specific information from them is challenging due to the lack of action labels.In this work, we introduce a novel framework, Any-point Trajectory Modeling (ATM), that utilizes video demonstrations by pre-training a trajectory model to predict future trajectories of arbitrary points within a video frame.Once trained, these trajectories provide detailed control guidance, enabling the learning of robust visuomotor policies with minimal action-labeled data.Across over 130 language-conditioned tasks we evaluated in both simulation and the real world, ATM outperforms strong video pre-training *First three authors contributed equally: Chuan Wen led the implementation and experiments.Xingyu Lin came up with the idea, supervised the technical development, and contributed to model debugging.John So implemented the Robot-to-robot transfer experiments and UniPi baselines.baselines by 80% on average.Furthermore, we show effective transfer learning of manipulation skills from human videos and videos from a different robot morphology.Visualizations and code are available at: https://xingyu-lin.github.io/atm. I. INTRODUCTIONComputer vision and natural language understanding have made significant advances in recent years [22,7], where the availability of large datasets plays a critical role.Similarly, in robotics, scaling up human demonstration data has been key for learning new skills [6,34,14], with a clear trend of performance improvement with larger datasets [29,6].However, human demonstrations, typically action-labeled trajectories collected via teleoperation devices [55,52], are time-consuming and labor-intensive to collect.For instance, collecting 130K trajectories in [6] took 17 months, making data collection a major bottleneck in robot learning.
Keywords
Related papers
Statistical Learning Theory
Yuhai Wu, Vladimir Vapnik
1999
Artificial intelligence: a modern approach
1995
Fractional Differential Equations
Igor Podlubný
2025
Applied Nonlinear Control
Jean-Jacques Slotine, Weiping Li
1991