IMPLEMENTASI SPEECH EMOTION RECOGNITION UNTUK KLASIFIKASI TINGKAT STRES DARI VOICE NOTE BERBASIS HYBRID CNN-LSTM

  • Muhamad Fakhri Khairil Imam
    Universitas Muhammadiyah Sukabumi
DOI: https://doi.org/10.23960/jitet.v14i3.10900
Keywords Speech Emotion Recognition, CNN-LSTM, Klasifikasi Tingkat Stres, Hybrid Fusion, Voice
Abstract Views (Last 12 Months)
6 Abstract Views
8 Downloads

Abstract

Kesehatan mental menjadi isu penting seiring meningkatnya tekanan psikologis akibat gaya hidup digital, namun deteksi stres masih mengandalkan instrumen self-report yang bersifat terjadwal dan subjektif. Penelitian ini mengembangkan sistem Speech Emotion Recognition (SER) berbasis voice note berbahasa Indonesia untuk mengklasifikasikan tingkat stres menggunakan pendekatan hybrid CNN-LSTM (akustik) dan Zero-Shot Classification IndoBERT (linguistik) dengan metodologi CRISP-DM. Model dilatih menggunakan dataset gabungan RAVDESS (1.440 file, pembagian speaker-independent) dan voice note primer berbahasa Indonesia (106 file) yang dipetakan ke tiga kategori Non_Stress, Mild_Stress, dan High_Stress, menghasilkan 1.546 sampel asli. Fitur akustik diekstraksi menggunakan 40 koefisien MFCC dengan panjang tetap 174 frame. Pengujian pada 256 sampel independen menghasilkan akurasi model CNN-LSTM akustik sebesar 62,11% dengan F1-score terbaik pada kelas High_Stress (0,69), sementara kelas Mild_Stress menunjukkan performa terendah (F1 0,41) akibat karakteristik akustik yang ambigu. Untuk mengatasi kesenjangan bahasa antara dataset RAVDESS dan voice note Indonesia, sistem menerapkan fusi hibrida dengan bobot linguistik 90% dan akustik 10%. Model diintegrasikan ke aplikasi web StressVoice berbasis Flask yang dilengkapi transkripsi Whisper AI, respons empatik, cetak surat rujukan PDF, dan integrasi konsultasi WhatsApp ke psikolog.

Mental health has become a critical issue amid rising psychological pressure from digital lifestyles, yet stress detection still relies on scheduled, subjective self-report instruments. This study develops a Speech Emotion Recognition (SER) system based on Indonesian voice notes to classify stress levels using a hybrid CNN-LSTM (acoustic) and IndoBERT Zero-Shot Classification (linguistic) approach under the CRISP-DM methodology. The model was trained on a combined RAVDESS dataset (1,440 files, speaker-independent split) and primary Indonesian voice notes (106 files) mapped into three categories, Non_Stress, Mild_Stress, and High_Stress, yielding 1,546 original samples. Acoustic features were extracted using 40 MFCC coefficients with a fixed length of 174 frames. Testing on 256 independent samples produced an acoustic CNN-LSTM accuracy of 62.11%, with the best F1-score on the High_Stress class (0.69), while Mild_Stress showed the lowest performance (F1 0.41) due to ambiguous acoustic characteristics. To bridge the language gap between RAVDESS and Indonesian voice notes, the system applies hybrid fusion weighting linguistic analysis at 90% and acoustic analysis at 10%. The model is integrated into the Flask-based StressVoice web application, equipped with Whisper AI transcription, empathetic responses.

Downloads

Download data is not yet available.

References

K. Mountzouris, I. Perikos, and I. Hatzilygeroudis, “Speech Emotion Recognition Using Convolutional Neural Networks with Attention Mechanism,” Electron., vol. 12, no. 20, pp. 1–31, 2023, doi: 10.3390/electronics12204376.

Vamsinath J, Varshini Bonagiri, Sandeep T, Meghana V, and Latha B, “Stress Detection Through Speech Analysis Using Machine Learning,” Int. J. Sci. Res. Sci. Technol., pp. 334–342, 2022, doi: 10.32628/ijsrst229437.

M. N. Aljufri and B. H. Prasetio, “Sistem Deteksi Tingkat Stress Menggunakan Suara dengan Metode Jaringan Saraf Tiruan dan Ekstraksi Fitur MFCC berbasis Raspberry Pi,” J. Pengemb. Teknol. Inf. dan Ilmu Komput., vol. 6, no. 11, pp. 5278–5285, 2022, [Online]. Available: https://j-ptiik.ub.ac.id/index.php/j-ptiik/article/view/11842

S. Latif, R. Rana, S. Khalifa, R. Jurdak, J. Qadir, and B. Schuller, “Survey of Deep Representation Learning for Speech Emotion Recognition,” IEEE Trans. Affect. Comput., vol. 14, no. 2, pp. 1634–1654, 2023, doi: 10.1109/TAFFC.2021.3114365.

M. E. P. Rismanto and I. Handayani, “Klasifikasi Emosi Berdasarkan Suara dengan Metode Convolutional Neural Network,” J. Inform. Univ. Pamulang, vol. 9, no. 4, pp. 163–171, 2025, doi: 10.32493/informatika.v9i4.45236.

Y. Hifny and A. Ali, “Efficient Arabic Emotion Recognition Using Deep Neural Networks,” ICASSP, IEEE Int. Conf. Acoust. Speech Signal Process. - Proc., vol. 2020-May, pp. 6710–6714, 2024, doi: 10.1109/ICASSP.2019.8683632.

H. Rheza Paleva and B. Henryranu Prasetio, “Penerapan Short Time Fourier Transform pada MFCC untuk Sistem Pengenalan Ucapan Tingkat Stres,” J. Pengemb. Teknol. Inf. dan Ilmu Komput., vol. 1, no. 1, pp. 1–9, 2024, [Online]. Available: http://j-ptiik.ub.ac.id

A. A. Kasim, M. Bakri, I. Mahmudi, R. Rahmawati, and Z. Zulnabil, “Artificial Intelligent for Human Emotion Detection with the Mel-Frequency Cepstral Coefficient (MFCC),” JUITA J. Inform., vol. 11, no. 1, p. 47, 2023, doi: 10.30595/juita.v11i1.15435.

F. KASYIDI, R. ILYAS, and N. M. ANNISA, “Peningkatan Kemampuan Pengenalan Emosi Melalui Suara dalam Bahasa Indonesia,” MIND J., vol. 6, no. 2, pp. 194–204, 2021, doi: 10.26760/mindjournal.v6i2.194-204.

N. N. Y. Truong et al., “DFAT: Dual-stage Fusion of Acoustic and Text feature for Speech Emotion Recognition,” Proc. 11th Int. Work. Vietnamese Lang. Speech Process., pp. 36–44, 2025, [Online]. Available: https://aclanthology.org/2025.vlsp-1.6/

T. Srivastaval, J. C. Chou, P. Shroffl, K. Livescu, and C. Graziul, “Speech Recognition For Analysis of Police Radio Communication,” Proc. 2024 IEEE Spok. Lang. Technol. Work. SLT 2024, pp. 906–912, 2024, doi: 10.1109/SLT61566.2024.10832157.

S. G. Tesfagergish, J. Kapočiūtė-Dzikienė, and R. Damaševičius, “Zero-Shot Emotion Detection for Semi-Supervised Sentiment Analysis Using Sentence Transformers and Ensemble Learning,” Appl. Sci., vol. 12, no. 17, 2022, doi: 10.3390/app12178662.

A. D. Fathulramdhan and B. H. Prasetyo, “Pengembangan Sistem Deteksi Stres Berbasis Suara Menggunakan Fitur MFCC, ZCR, dan SC Dengan Metode Artificial Neural Network (ANN),” J. Pengemb. Teknol. Inf. dan Ilmu Komput., vol. 9, no. 9, pp. 2548–964, 2025, [Online]. Available: http://j-ptiik.ub.ac.id

M. N. Hadi and R. W. Sari, “Multi-Label Emotion Detection for Mental Health Monitoring Using Deep CNN and Visual Attention,” J. Artif. Intell. Softw. Eng., vol. 5, no. 2, pp. 597–605, 2025, doi: 10.30811/jaise.v5i2.6961.

A. Roihan, R. Zein, M. R. Aprianti, F. N. Izzah, and A. Luthfunnisa, “Analisis Teknologi Speech Emotion Recognition (SER): Pendekatan Fitur Akustik, Klasifikasi, Keamanan Dan Implementasi Pada Sistem Portabel,” J. Sist. Inf. dan Teknol., vol. 5, no. 2, pp. 144–149, 2025, doi: 10.56995/sintek.v5i2.167.

A. S. Sams and A. Zahra, “Multimodal music emotion recognition in Indonesian songs based on CNN-LSTM, XLNet transformers,” Bull. Electr. Eng. Informatics, vol. 12, no. 1, pp. 355–364, 2023, doi: 10.11591/eei.v12i1.4231.

F. Makhmudov, A. Kutlimuratov, and Y. I. Cho, “Hybrid LSTM–Attention and CNN Model for Enhanced Speech Emotion Recognition,” Appl. Sci., vol. 14, no. 23, 2024, doi: 10.3390/app142311342.

A. Rianti, N. W. A. Majid, and A. Fauzi, “CRISP-DM: Metodologi Proyek Data Science,” Pros. Semin. Nas. Teknol. Inf. dan Bisnis 2023, pp. 107–104, 2023.

G. T. Waleed and S. H. Shaker, “Speech Emotion Recognition on MELD and RAVDESS Datasets Using CNN,” Inf., vol. 16, no. 7, 2025, doi: 10.3390/info16070518.

R. Ullah et al., “Speech Emotion Recognition Using Convolution Neural Networks and Multi-Head Convolutional Transformer,” Sensors, vol. 23, no. 13, pp. 1–20, 2023, doi: 10.3390/s23136212.

R. Akyas Hifdzi Rahman, A. Adi Sunarto, and A. Asriyanik, “Penerapan You Only Look Once (Yolo) V8 Untuk Deteksi Tingkat Kematangan Buah Manggis,” JATI (Jurnal Mhs. Tek. Inform., vol. 8, no. 5, pp. 10566–10571, 2024, doi: 10.36040/jati.v8i5.10979.

T. Putra, Dzaky, Putratama , Al Fikri, “Speech Emotion Recognition Deep Learning Karakter Kucing Interaktif Di Game Unity,” JITET (Jurnal Inform. dan Tek. Elektro Ter., vol. 14, no. 1, pp. 1583–1591, 2026.

Cover
Published
2026-08-13
How to Cite
Khairil Imam, M. F. (2026). IMPLEMENTASI SPEECH EMOTION RECOGNITION UNTUK KLASIFIKASI TINGKAT STRES DARI VOICE NOTE BERBASIS HYBRID CNN-LSTM. Jurnal Informatika Dan Teknik Elektro Terapan, 14(3). https://doi.org/10.23960/jitet.v14i3.10900