首页 /研究 /Learn-Gen-Plan: Bridging the Gap Between Vision Language Models and Real-World Long-Horizon Dexterous Manipulations
MANIPULATION

Learn-Gen-Plan: Bridging the Gap Between Vision Language Models and Real-World Long-Horizon Dexterous Manipulations

Hao Peng, Shaowei Cui, Junhang Wei, Tao Lu, Yinghao Cai, Shuo Wang

发表年份
2025
引用次数
5

摘要

Long-horizon dexterous tasks have been a long-standing problem in robotic manipulation. Previous studies have developed task and motion planning, imitation learning, and reinforcement learning methods for long-horizon manipulations. However, these methods are hard to achieve efficient planning for new tasks. Empowered with the Vision Language Model (VLM), recent studies significantly improve the generalization of robot systems. However, these works are only verified in simple pick-and-place tasks due to limited skills. To this end, we propose the Learn-Gen-Plan (LGP), which combines the VLM and learning-based primitives to endow robots with the ability to efficiently plan and complete various long-horizon dexterous tasks. LGP contains two key phases: skill generation and task planning. In skill generation, the Skill Generator is proposed to utilize the learned key primitives and hand-crafted trivial primitives to generate adaptive robot skills. In task planning, the Multimodal Planner generates the robot plan based on image observation, generated skills, and text prompts. We set up a series of dexterous tasks (e.g., cable routing, peg-in-hole assembly) in a real-world lighting circuit wiring scenario to evaluate LGP. The experimental results show that LGP efficiently generates robot plans with learned skills, controlling the robot to complete various multi-step cable wiring tasks.

关键词

Bridging (networking)Plan (archaeology)HorizonComputer scienceArtificial intelligenceComputer visionMachine visionEngineeringHuman–computer interactionMathematics

相关论文

查看 MANIPULATION 分类全部论文