Нейромережевий метод розпізнавання емоційного стану мовлення в системах контакт-центрів на основі архітектури CNN-BiLSTM з модифікованим механізмом уваги
Вантажиться...
Файли
Дата
Автори
Назва журналу
Номер ISSN
Назва тому
Анотація
The relevance of this research lies in the need to improve the efficiency of decision-support systems in contact centers for automated analysis of the emotional states of operators and clients. Detecting emotional tension in the voice enables timely adjustments to interactions, enhancing service quality and operator performance. Traditional audio signal processing methods based on Mel-Frequency Cepstral Coefficients (MFCCs) have limitations in preserving complete acoustic information, thereby reducing the accuracy of emotion recognition. This work proposes a neural network method that combines Convolutional Neural Networks (CNNs) and Bidirectional Long Short-Term Memory (BiLSTM) networks with a modified Attention mechanism. The first stage involves loading a Ukrainian-language audio dataset and performing preliminary data processing, including amplitude normalization, noise filtering, speech segmentation, conversion to Mel-spectrograms, and extraction of low-level descriptors (LLD), such as energy and fundamental frequency (F0). The input data are formed into fixed-dimension tensors for neural network analysis. During the feature extraction stage, the CNN automatically identifies local spectral characteristics of the signal, including intensity, frequency components, and intonational peaks. Each convolutional block is complemented with batch normalization to stabilize training and accelerate convergence. To model the temporal dynamics of the emotional state, bidirectional BiLSTM layers are applied, taking into account the context of preceding and subsequent signal segments. The Attention mechanism forms a context vector as a weighted sum of features, determining the relative importance of individual time frames, and then passes it to a fully connected layer with a tanh activation function. In this work, the modified Attention mechanism refers to the integration of audio recording metadata (signal duration, 640 kbps bitrate, technical identifiers) into the formation of the context vector via a Multi-Weighting System (MWS). It allows simultaneous consideration of local spectral features, temporal dynamics of the audio signal, and the relevance of individual speech segments. The scientific novelty lies in the development of a neural network method that combines CNN–BiLSTM–Attention with a multimodal weighting mechanism, integrating audio metadata into the formation of the context vector. This architecture provides increased accuracy in recognizing emotional tension. A comparative analysis of traditional MFCC and LLD was conducted, demonstrating the advantage of LLD: the baseline CNN accuracy increased 82,42 % to 9,.00 %, and integrating the Attention layer further improved accuracy by 1.5...2%. The highest achieved accuracy was 93,48%.
Опис
УДК
Тип документа
Мова
ISSN
Бібліографічний опис
Овчаренко М. А., Каштан В. Ю. Нейромережевий метод розпізнавання емоційного стану мовлення в системах контакт-центрів на основі архітектури CNN-BiLSTM з модифікованим механізмом уваги // Вісник Вінницького політехнічного інституту. 2026. № 3. С. 52-60. URI: https://visnyk.vntu.edu.ua/index.php/visnyk/article/view/3514.
Схвалення
Рецензія
Доповнено
Цитується в
Список використаної літератури (11)
- N. Dhanpat, F. D. Modau, P. Lugisani, R. Mabojane, and M. Phiri, “Exploring employee retention and intention to leave within a call center,” SA J. Hum. Resour. Manag., vol. 16, pp. 1-13, 2018.[Online]. Available: https://www.researchgate.net/publication/323869688_Exploring_employee_retention_and_intention_to_leave_within_a_call_centre .
- B. Karakus, and G. Aydin, “Call center performance evaluation using big data analytics,” in Proc. 2016 International Symposium on Networks, Computers and Communications (ISNCC), Hammamet, Tunisia, 2016, pp. 1-6 . https://doi.org/10.1109/ISNCC.2016.7746116.
- J. Cho, R. Pappagari, P. Kulkarni, J. Villalba, Y. Carmiel, and N. Dehak, “Deep neural networks for emotion recognition combining audio and transcripts,” arXiv preprint arXiv:1911.00432, 2019. https://doi.org/10.21437/Interspeech.2018-2466.
- X. Liu, J. van de Weijer, and A. D. Bagdanov, “Exploiting unlabeled data in cnns by self-supervised learning to rank,” IEEE Trans. Pattern Anal. Mach. Intell., vol. 41, pp. 1862-1878, 2019, https://doi.org/10.48550/arXiv.1902.06285.
- F. Karim, S. Majumdar, and H. Darabi, “Insights into LSTM fully convolutional networks for time series classification,” IEEE Access, vol. 7, pp. 67718-67725, 2019, https://doi.org/10.1109/ACCESS.2019.2916828.
- N. Shabrina, F. Kasyidi, and R. Ilyas, “A BiLSTM-Based Approach for Speech Emotion Recognition in Conversational Indonesian Audio using SMOTE,” Jurnal Teknik Informatika (Jutif), vol. 6, pp. 3173-3187, 2025, https://doi.org/10.52436/1.jutif.2025.6.5.5183.
- S. Ayadi, and Z. Lachiri, “A combined CNN-LSTM Network for Audio Emotion Recognition using Speech and Song at-tributs,” in Proc. 2022 6th International Conference on Advanced Technologies for Signal and Image Processing (ATSIP), Sfax, Tunisia, 2022, pp. 1-6, https://doi.org/10.1109/ATSIP55956.2022.9805924.
- A. Ahmed, S. Toral, K. Shaalan, and Y. Hifny, “Agent Productivity Modeling in a Call Center Domain Using Attentive Convolutional Neural Networks,” Sensors, vol. 20, p. 5489, 2020, https://doi.org/10.3390/s20195489.
- A. Norouzian, B. Mazoure, D. Connolly, and D. Willett, “Exploring attention mechanism for acoustic-based classification of speech utterances into system-directed and non-system-directed,” in Proc. ICASSP 2019-2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Brighton, UK, 2019, pp. 7310-7314, https://doi.org/10.48550/arXiv.1902.00570.
- L. Shu, J. Xie, M. Yang, Z. Li, Z. Li, D. Liao, X. Xu, and X. Yang, “A review of emotion recognition using physiologi-cal signals,” Sensors, vol. 18, p. 2074, 2018, https://doi.org/10.3390/s18072074.