首页 /研究 /Learning to Plan Optimistically: Uncertainty-Guided Deep Exploration via Latent Model Ensembles

OTHER

Learning to Plan Optimistically: Uncertainty-Guided Deep Exploration via Latent Model Ensembles

Tim Seyde, Wilko Schwarting, Sertac Karaman, Daniela Rus

发表年份: 2020
访问权限: 开放获取

摘要

Learning complex robot behaviors through interaction requires structured exploration. Planning should target interactions with the potential to optimize long-term performance, while only reducing uncertainty where conducive to this objective. This paper presents Latent Optimistic Value Exploration (LOVE), a strategy that enables deep exploration through optimism in the face of uncertain long-term rewards. We combine latent world models with value function estimation to predict infinite-horizon returns and recover associated uncertainty via ensembling. The policy is then trained on an upper confidence bound (UCB) objective to identify and select the interactions most promising to improve long-term performance. We apply LOVE to visual robot control tasks in continuous action spaces and demonstrate on average more than 20% improved sample efficiency in comparison to state-of-the-art and other exploration objectives. In sparse and hard to explore environments we achieve an average improvement of over 30%.

关键词

cs.LGcs.AIcs.RO

Learning to Plan Optimistically: Uncertainty-Guided Deep Exploration via Latent Model Ensembles

摘要

关键词

相关论文

一种面向线弧增材制造的电动汽车结构可制造性拓扑优化的双环框架

几何数字孪生：一种用于航空发动机装配精度预测的数字智能模型

通过人工智能驱动的机器人技术革新产业

新型大口径偏置馈电可展开天线设计与动态性能预测