首页 /研究 /Embodied BERT: A Transformer Model for Embodied, Language-guided Visual\n Task Completion
OTHER

Embodied BERT: A Transformer Model for Embodied, Language-guided Visual\n Task Completion

Alessandro Suglia, Qiaozi Gao, Jesse Thomason, Govind Thattai, Gaurav S. Sukhatme

发表年份
2021
引用次数
30
访问权限
开放获取

摘要

Language-guided robots performing home and office tasks must navigate in and\ninteract with the world. Grounding language instructions against visual\nobservations and actions to take in an environment is an open challenge. We\npresent Embodied BERT (EmBERT), a transformer-based model which can attend to\nhigh-dimensional, multi-modal inputs across long temporal horizons for\nlanguage-conditioned task completion. Additionally, we bridge the gap between\nsuccessful object-centric navigation models used for non-interactive agents and\nthe language-guided visual task completion benchmark, ALFRED, by introducing\nobject navigation targets for EmBERT training. We achieve competitive\nperformance on the ALFRED benchmark, and EmBERT marks the first\ntransformer-based model to successfully handle the long-horizon, dense,\nmulti-modal histories of ALFRED, and the first ALFRED model to utilize\nobject-centric navigation targets.\n

关键词

Embodied cognitionTransformerComputer scienceRobotTask (project management)Language understandingBenchmark (surveying)Artificial intelligenceLanguage modelBridge (graph theory)

相关论文

查看 OTHER 分类全部论文