Home /Research /The World According to a Social Robot -- Augmenting Human-Robot Dialogue With Vision Language Models
HRI

The World According to a Social Robot -- Augmenting Human-Robot Dialogue With Vision Language Models

Thomas Sievers

Year
2026
Access
Open access

Abstract

Vision Language Models (VLMs) enable robots to visually perceive their environment as well as the actions and characteristics of their conversation partner or humans in collaboration. Especially for social robots deployed in everyday settings and for uncomplicated, natural use, it is essential that the robot has an understanding of situations that is appropriate to human customs. This paper presents initial experiences with the application of a Mistral AI language model with a Pepper robot for Human-Robot Interaction (HRI) in dialogue, as well as an investigation of the effects of additional visual information on response time in different models. The results show that incorporating visual information adds context to the dialogue with only a moderate increase in response time, enabling both the robot and the human to take into account unspoken elements of the situation. Furthermore, using an LLM hosted in Europe offers a solution that complies with European data protection regulations and can therefore facilitate real-life applications more easily.

Keywords

cs.HCcs.RO

Related papers

Browse all HRI papers