DexVIP: Learning Dexterous Grasping with Human Hand Pose Priors from\n Video
Priyanka Mandikal, Kristen Grauman
- 发表年份
- 2022
- 引用次数
- 12
- 访问权限
- 开放获取
摘要
Dexterous multi-fingered robotic hands have a formidable action space, yet\ntheir morphological similarity to the human hand holds immense potential to\naccelerate robot learning. We propose DexVIP, an approach to learn dexterous\nrobotic grasping from human-object interactions present in in-the-wild YouTube\nvideos. We do this by curating grasp images from human-object interaction\nvideos and imposing a prior over the agent's hand pose when learning to grasp\nwith deep reinforcement learning. A key advantage of our method is that the\nlearned policy is able to leverage free-form in-the-wild visual data. As a\nresult, it can easily scale to new objects, and it sidesteps the standard\npractice of collecting human demonstrations in a lab -- a much more expensive\nand indirect way to capture human expertise. Through experiments on 27 objects\nwith a 30-DoF simulated robot hand, we demonstrate that DexVIP compares\nfavorably to existing approaches that lack a hand pose prior or rely on\nspecialized tele-operation equipment to obtain human demonstrations, while also\nbeing faster to train. Project page:\nhttps://vision.cs.utexas.edu/projects/dexvip-dexterous-grasp-pose-prior\n
关键词
相关论文
Statistical Learning Theory
Yuhai Wu, Vladimir Vapnik
1999
Artificial intelligence: a modern approach
1995
Applied Nonlinear Control
Jean-Jacques Slotine, Weiping Li
1991
A new optimizer using particle swarm theory
R.C. Eberhart, James Kennedy
2002