Home /Research /Audio-Visual Conversation Analysis by Smart Posterboard and Humanoid Robot
HRI

Audio-Visual Conversation Analysis by Smart Posterboard and Humanoid Robot

Tatsuya Kawahara, Koji Inoue, Divesh Lala, Katsuya Takanashi

Year
2018
Citations
2

Abstract

This paper addresses audio-visual signal processing for conversation analysis, which involves multi-modal behavior detection and mental-state recognition. We have investigated prediction of turn-taking by the audience in a poster session from their multi-modal behaviors, and found out that the eye-gaze provides an important cue compared with head nodding and verbal backchannels. This finding has been applied to audio-visual speaker diarization by combining eye-gaze information. We are now investigating engagement recognition in human-robot interaction based on the same scheme. Robust and realtime detection of laughing, backchannels and nodding is realized based on LSTM-CTC. We introduce a latent “character” model to cope with the subjectivity and variations of engagement annotations. Experimental evaluations demonstrate that (1) the latent character model is effective, (2) automatic behavior detection is robust and does not degrade the engagement recognition accuracy, and (3) the eye-gaze is the most important feature among others.

Keywords

ConversationHumanoid robotComputer scienceHuman–computer interactionAudio visualRobotMultimediaSpeech recognitionComputer visionArtificial intelligence

Related papers

Browse all HRI papers