首页 /研究 /Deep Neural Network for visual Emotion Recognition based on ResNet50 using Song-Speech characteristics
LEARNING

Deep Neural Network for visual Emotion Recognition based on ResNet50 using Song-Speech characteristics

Souha Ayadi, Zied Lachiri

发表年份
2022
引用次数
11

摘要

Visual emotion recognition is a very large field. It plays a very important role in different domains such as security, robotics, and medical tasks. The visual tasks could be either image or video. Unlike the image processing, the difficulty of video processing is always a challenge due to changes in information over time variation. Significant performance improvements when applying deep learning algorithms to video processing. This paper presents a deep neural network based on ResNet50 model. The latter is conducted on the Ryerson Audio-Visual Database of Emotional Speech and Song (RAVDESS) due to the variance of the nature of the data exists which is speech and song. The choice of ResNet model is based on the ability of facing different problems such as of vanishing gradients, the performing stability offered by this model, the ability of CNN for feature extraction which is considered to be the base architecture for ResNet, and the ability of improving the accuracy results and minimizing the loss. The achieved results are 57.73% for song and 55.52% for speech. Results shows that the Resnet50 model is suitable for both speech and song while maintaining performance stability.

关键词

Computer scienceSpeech recognitionArtificial intelligenceFeature extractionArtificial neural networkField (mathematics)Deep learningSpeech processingStability (learning theory)Feature (linguistics)

相关论文

查看 LEARNING 分类全部论文