High Acceleration Reinforcement Learning for Real-World Juggling with\n Binary Rewards
Kai Ploeger, Michael Lutter, Jan Peters
- 发表年份
- 2020
- 引用次数
- 13
- 访问权限
- 开放获取
摘要
Robots that can learn in the physical world will be important to en-able\nrobots to escape their stiff and pre-programmed movements. For dynamic\nhigh-acceleration tasks, such as juggling, learning in the real-world is\nparticularly challenging as one must push the limits of the robot and its\nactuation without harming the system, amplifying the necessity of sample\nefficiency and safety for robot learning algorithms. In contrast to prior work\nwhich mainly focuses on the learning algorithm, we propose a learning system,\nthat directly incorporates these requirements in the design of the policy\nrepresentation, initialization, and optimization. We demonstrate that this\nsystem enables the high-speed Barrett WAM manipulator to learn juggling two\nballs from 56 minutes of experience with a binary reward signal. The final\npolicy juggles continuously for up to 33 minutes or about 4500 repeated\ncatches. The videos documenting the learning process and the evaluation can be\nfound at https://sites.google.com/view/jugglingbot\n
关键词
相关论文
Statistical Learning Theory
Yuhai Wu, Vladimir Vapnik
1999
Artificial intelligence: a modern approach
1995
Fractional Differential Equations
Igor Podlubný
2025
Applied Nonlinear Control
Jean-Jacques Slotine, Weiping Li
1991