<link rel="stylesheet" href="styles.f3b1fba60ec7970c.css">

Нейромережевий метод розпізнавання емоційного стану мовлення в системах контакт-центрів на основі архітектури CNN-BiLSTM з модифікованим механізмом уваги

dc.contributor.authorOvcharenko, M. A.en
dc.contributor.authorKashtan, V. Yu.en
dc.date.accessioned2026-08-20T09:33:12Z
dc.date.available2026-08-20T09:33:12Z
dc.date.issued2026
dc.description.abstractThe relevance of this research lies in the need to improve the efficiency of decision-support systems in contact centers for automated analysis of the emotional states of operators and clients. Detecting emotional tension in the voice enables timely adjustments to interactions, enhancing service quality and operator performance. Traditional audio signal processing methods based on Mel-Frequency Cepstral Coefficients (MFCCs) have limitations in preserving complete acoustic information, thereby reducing the accuracy of emotion recognition. This work proposes a neural network method that combines Convolutional Neural Networks (CNNs) and Bidirectional Long Short-Term Memory (BiLSTM) networks with a modified Attention mechanism. The first stage involves loading a Ukrainian-language audio dataset and performing preliminary data processing, including amplitude normalization, noise filtering, speech segmentation, conversion to Mel-spectrograms, and extraction of low-level descriptors (LLD), such as energy and fundamental frequency (F0). The input data are formed into fixed-dimension tensors for neural network analysis. During the feature extraction stage, the CNN automatically identifies local spectral characteristics of the signal, including intensity, frequency components, and intonational peaks. Each convolutional block is complemented with batch normalization to stabilize training and accelerate convergence. To model the temporal dynamics of the emotional state, bidirectional BiLSTM layers are applied, taking into account the context of preceding and subsequent signal segments. The Attention mechanism forms a context vector as a weighted sum of features, determining the relative importance of individual time frames, and then passes it to a fully connected layer with a tanh activation function. In this work, the modified Attention mechanism refers to the integration of audio recording metadata (signal duration, 640 kbps bitrate, technical identifiers) into the formation of the context vector via a Multi-Weighting System (MWS). It allows simultaneous consideration of local spectral features, temporal dynamics of the audio signal, and the relevance of individual speech segments. The scientific novelty lies in the development of a neural network method that combines CNN–BiLSTM–Attention with a multimodal weighting mechanism, integrating audio metadata into the formation of the context vector. This architecture provides increased accuracy in recognizing emotional tension. A comparative analysis of traditional MFCC and LLD was conducted, demonstrating the advantage of LLD: the baseline CNN accuracy increased 82,42&nbsp;% to 9,.00&nbsp;%, and integrating the Attention layer further improved accuracy by 1.5...2%. The highest achieved accuracy was 93,48%.en
dc.identifier.citationОвчаренко М. А., Каштан В. Ю. Нейромережевий метод розпізнавання емоційного стану мовлення в системах контакт-центрів на основі архітектури CNN-BiLSTM з модифікованим механізмом уваги // Вісник Вінницького політехнічного інституту. 2026. № 3. С. 52-60. URI: https://visnyk.vntu.edu.ua/index.php/visnyk/article/view/3514.uk
dc.identifier.doihttps://doi.org/10.31649/1997-9266-2026-186-3-52-60
dc.identifier.issn1997-9274
dc.identifier.orcidhttps://orcid.org/0009-0006-0730-0913
dc.identifier.orcidhttps://orcid.org/0000-0002-0395-5895
dc.identifier.udc004.9
dc.identifier.urihttps://ir.lib.vntu.edu.ua/handle/123456789/52342
dc.language.isouk_UAuk_UA
dc.publisherВінницький національний технічний університетuk
dc.relation.ispartofВісник Вінницького політехнічного інституту. № 3 : 52-60.uk
dc.relation.referencesN. Dhanpat, F. D. Modau, P. Lugisani, R. Mabojane, and M. Phiri, “Exploring employee retention and intention to leave within a call center,” SA J. Hum. Resour. Manag., vol. 16, pp. 1-13, 2018.[Online]. Available: https://www.researchgate.net/publication/323869688_Exploring_employee_retention_and_intention_to_leave_within_a_call_centre .en
dc.relation.referencesB. Karakus, and G. Aydin, “Call center performance evaluation using big data analytics,” in Proc. 2016 International Symposium on Networks, Computers and Communications (ISNCC), Hammamet, Tunisia, 2016, pp. 1-6 . https://doi.org/10.1109/ISNCC.2016.7746116.en
dc.relation.referencesJ. Cho, R. Pappagari, P. Kulkarni, J. Villalba, Y. Carmiel, and N. Dehak, “Deep neural networks for emotion recognition combining audio and transcripts,” arXiv preprint arXiv:1911.00432, 2019. https://doi.org/10.21437/Interspeech.2018-2466.en
dc.relation.referencesX. Liu, J. van de Weijer, and A. D. Bagdanov, “Exploiting unlabeled data in cnns by self-supervised learning to rank,” IEEE Trans. Pattern Anal. Mach. Intell., vol. 41, pp. 1862-1878, 2019, https://doi.org/10.48550/arXiv.1902.06285.en
dc.relation.referencesF. Karim, S. Majumdar, and H. Darabi, “Insights into LSTM fully convolutional networks for time series classification,” IEEE Access, vol. 7, pp. 67718-67725, 2019, https://doi.org/10.1109/ACCESS.2019.2916828.en
dc.relation.referencesN. Shabrina, F. Kasyidi, and R. Ilyas, “A BiLSTM-Based Approach for Speech Emotion Recognition in Conversational Indonesian Audio using SMOTE,” Jurnal Teknik Informatika (Jutif), vol. 6, pp. 3173-3187, 2025, https://doi.org/10.52436/1.jutif.2025.6.5.5183.en
dc.relation.referencesS. Ayadi, and Z. Lachiri, “A combined CNN-LSTM Network for Audio Emotion Recognition using Speech and Song at-tributs,” in Proc. 2022 6th International Conference on Advanced Technologies for Signal and Image Processing (ATSIP), Sfax, Tunisia, 2022, pp. 1-6, https://doi.org/10.1109/ATSIP55956.2022.9805924.en
dc.relation.referencesA. Ahmed, S. Toral, K. Shaalan, and Y. Hifny, “Agent Productivity Modeling in a Call Center Domain Using Attentive Convolutional Neural Networks,” Sensors, vol. 20, p. 5489, 2020, https://doi.org/10.3390/s20195489.en
dc.relation.referencesA. Norouzian, B. Mazoure, D. Connolly, and D. Willett, “Exploring attention mechanism for acoustic-based classification of speech utterances into system-directed and non-system-directed,” in Proc. ICASSP 2019-2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Brighton, UK, 2019, pp. 7310-7314, https://doi.org/10.48550/arXiv.1902.00570.en
dc.relation.referencesL. Shu, J. Xie, M. Yang, Z. Li, Z. Li, D. Liao, X. Xu, and X. Yang, “A review of emotion recognition using physiologi-cal signals,” Sensors, vol. 18, p. 2074, 2018, https://doi.org/10.3390/s18072074.en
dc.relation.referencesT. Anvarjon, Mustaqeem, and S. Kwon, “Deep-Net: A Lightweight CNN-Based Speech Emotion Recognition System Using Deep Frequency Features,” Sensors, vol. 20, p. 5212, 2020, https://doi.org/10.3390/s20185212.en
dc.relation.urihttps://visnyk.vntu.edu.ua/index.php/visnyk/article/view/3514
dc.subjectspeech emotion recognitionen
dc.subjectconvolutional neural networksen
dc.subjectbidirectional long-term memory networksen
dc.subjectmultimodal weightingen
dc.subjectaudio analysisen
dc.titleНейромережевий метод розпізнавання емоційного стану мовлення в системах контакт-центрів на основі архітектури CNN-BiLSTM з модифікованим механізмом увагиuk
dc.title.alternativeNeural Network Method for Emotion Speech State Recognition in Contact Centre Systems Based on CNN-BiLSTM Architecture with a Modified Attention Mechanismen
dc.typeArticle, professional native edition
dc.typeArticle

Файли

Контейнер файлів

Зараз показуємо 1 - 1 з 1
Вантажиться...
Ескіз
Назва:
206349.pdf
Розмір:
53,88 KB
Формат:
Adobe Portable Document Format

Ліцензійна угода

Зараз показуємо 1 - 1 з 1
Вантажиться...
Ескіз
Назва:
license.txt
Розмір:
129 B
Формат:
Plain Text
Опис: