Multi-channel Convolution Classroom Speech Emotion Recognition Model Based on Improved Attention Mechanism
LIANG Kejin
ZHANG Haijun
Abstract:Aiming at the situation that increasing the depth and width of the network does not significantly improve the recogni-tion accuracy in speech emotion recognition research,the dual-channel attention mechanism is improved.By combining the chan-nel attention mechanism and the spatial attention mechanism,the spatial attention mechanism is combined.The convolution part is improved to a two-layer convolution in order to extract more valuable contextual semantic information.For a single emotional feature cannot effectively represent speech emotion,multiple single emotional features are fused to increase the emotional representation ability of the feature.The model obtained a recognition accuracy of 85.24%under the Chinese Affective Database of the Institute of Automation,Chinese Academy of Sciences(CASIA),and a recognition accuracy of 86.58%on the Emo-DB dataset,proving the effectiveness of the model.For the real classroom speech data,recall rate,F1 value and accuracy rate of the model in the experi-ment reached 77.77%,80.76%,and 79.24%,respectively,reflecting good practicability.
Keywords:emotion recognitiondeep learningspeech emotion recognitionneural networkunbalanced dataset
Publication Date:2024-09-20
Online Publishing Date:2025-08-15(First online date of this platform, not the publication date of the document)
Pages:6( 2645-2650 )
