Home /Research /Making Sense of Vision and Touch: Learning Multimodal Representations\n for Contact-Rich Tasks
HRI

Making Sense of Vision and Touch: Learning Multimodal Representations\n for Contact-Rich Tasks

Michelle A. Lee, Yuke Zhu, Peter Zachares, Matthew Tan, Silvio Savarese, Li Fei-Fei, Animesh Garg, Jeannette Bohg

Year
2019
Citations
2
Access
Open access

Abstract

Contact-rich manipulation tasks in unstructured environments often require\nboth haptic and visual feedback. It is non-trivial to manually design a robot\ncontroller that combines these modalities which have very different\ncharacteristics. While deep reinforcement learning has shown success in\nlearning control policies for high-dimensional inputs, these algorithms are\ngenerally intractable to deploy on real robots due to sample complexity. In\nthis work, we use self-supervision to learn a compact and multimodal\nrepresentation of our sensory inputs, which can then be used to improve the\nsample efficiency of our policy learning. Evaluating our method on a peg\ninsertion task, we show that it generalizes over varying geometries,\nconfigurations, and clearances, while being robust to external perturbations.\nWe also systematically study different self-supervised learning objectives and\nrepresentation learning architectures. Results are presented in simulation and\non a physical robot.\n

Keywords

Reinforcement learningComputer scienceHaptic technologyRepresentation (politics)Artificial intelligenceModalitiesTask (project management)RobotHuman–computer interactionController (irrigation)

Related papers

Browse all HRI papers