Learning Visual-Audio Representations for Voice-Controlled Robots

Peixin Chang, Shuijing Liu, Katherine Driggs-Campbell

发表年份: 2021
访问权限: 开放获取

摘要

Inspired by sensorimotor theory, we propose a novel pipeline for task-oriented voice-controlled robots. Previous method relies on a large amount of labels as well as task-specific reward functions. Not only can such an approach hardly be improved after the deployment, but also has limited generalization across robotic platforms and tasks. To address these problems, we learn a visual-audio representation (VAR) that associates images and sound commands with minimal supervision. Using this representation, we generate an intrinsic reward function to learn robot policies with reinforcement learning, which eliminates the laborious reward engineering process. We demonstrate our approach on various robotic platforms, where the robots hear an audio command, identify the associated target object, and perform precise control to fulfill the sound command. We show that our method outperforms previous work across various sound types and robotic tasks even with fewer amount of labels. We successfully deploy the policy learned in a simulator to a real Kinova Gen3. We also demonstrate that our VAR and the intrinsic reward function allows the robot to improve itself using only a small amount of labeled data collected in the real world.

关键词

cs.ROcs.AI

Learning Visual-Audio Representations for Voice-Controlled Robots

摘要

关键词

相关论文

面向学习与规划的并行可微可达性：具有认证神经动力学与控制器的系统

人工智能增强的智能焊接岛：基础模型革新制造业

基于深度强化学习和动态图神经网络的多任务机器人调度代理

基于微调与AAS增强检索的LLM驱动自动化DFA评估