ISR-FABEL: A Landmark-Based Dataset and Multimodal Emotion Recognition Framework for Child-Robot Interaction
Beatriz A. Ribeiro, Ana Lopes, Carlos Carona, Urbano Nunes
- Year
- 2025
- Citations
- 3
Abstract
Emotion recognition plays a crucial role in enhancing human-robot interaction (HRI), particularly in contexts involving children where understanding emotional states can improve engagement and therapeutic interventions. Current approaches primarily focus on facial expressions, overlooking the valuable emotional insights provided by body posture. The present study introduces a multimodal framework that integrates both facial and body posture landmarks to capture a more comprehensive representation of emotional expression. This work introduces ISR-FABEL, a new dataset collected from child-robot interactions involving 88 children, including both real-world sessions with the NAO robot and controlled emotion imitation tasks. The proposed architecture incorporates a Gated Recurrent Unit (GRU)-based model for facial emotion recognition, a Random Forest classifier leveraging engineered features from body posture, and a late-fusion strategy to combine both modalities. Results show that the multimodal approach achieves 72.53% classification accuracy, significantly outperforming unimodal models (face-only: 61.54%, body-only: 39.47%), underscoring the complementary nature of facial and bodily cues in emotion recognition. By relying on 2D and 3D landmark data rather than raw visual input, the framework also supports practical deployment in sensitive environments. The resulting dataset and baseline model aim to support future research in affective computing and child-robot interaction.
Keywords
Related papers
Probabilistic graphical models : principles and techniques
Daniel L. Koller, Nir Friedman
2009
Color indexing
Michael J. Swain, Dana H. Ballard
1991
Intelligence without representation
Rodney A. Brooks
1991
Simultaneous localization and mapping: part I
Hugh Durrant‐Whyte, T. Bailey
2006