audio-visual emotion recognition based on a deep convolutional neural network

Fa | Ar | En

audio-visual emotion recognition based on a deep convolutional neural network


نویسنده	aghajani khadijeh
منبع	journal of ai and data mining - 2022 - دوره : 10 - شماره : 4 - صفحه:529 -537
چکیده	Emotion recognition has several applications in various fields, including human-computer interactions. in the recent years, various methods have been proposed to recognize emotion using facial or speech information, while the fusion of these two has been paid less attention in emotion recognition. in this work, first of all, the use of only face or speech information in emotion recognition is examined. for emotion recognition through speech, a pre-trained network called yamnet is used to extract the features. after passing through a convolutional neural network (cnn), the extracted features are then fed into a bi-lstm with an attention mechanism to perform the recognition. for emotion recognition through facial information, a deep cnn-based model is proposed. finally, after reviewing these two approaches, an emotion detection framework based on the fusion of these two models is proposed. the ryerson audio-visual database of emotional speech and song (ravdess) containing videos taken from 24 actors (12 men and 12 women) with 8 categories is used to evaluate the proposed model. the results of the implementation show that a combination of the face and speech information improves the performance of the emotion recognizer.
کلیدواژه	speech emotion recognition ,facial emotion recognition ,deep learning ,transfer learning
آدرس	university of mazandaran, department of computer engineering, iran
پست الکترونیکی	kh.aghajani@umz.ac.ir



Authors