Perbandingan Performa Model Sistem Pengenalan Suara Tradisional Dan Modern Pada Bahasa Indonesia
Abstract
Sistem pengenalan suara merupakan komponen penting pada robot untuk mendukung komunikasi alami antara manusia dan robot. Keterbatasan dataset berbahasa Indonesia dapat memengaruhi kinerja model pengenalan suara, sehingga diperlukan analisis untuk menentukan model yang sesuai. Penelitian ini bertujuan membandingkan performa empat model, yaitu Hidden Markov Model (HMM), HMM-Gaussian Mixture Model (HMM-GMM), HMM-GMM-Deep Neural Network (HMM-GMM-DNN), dan Transformer Vanilla menggunakan dataset suara bahasa Indonesia yang dikumpulkan secara mandiri. Dataset terdiri atas 1.250 data audio dari 50 kalimat, yang meliputi 25 kalimat tanya dan 25 kalimat perintah dengan 25 kali pengulangan. Model HMM, HMM-GMM, dan HMM-GMM-DNN dilatih menggunakan Kaldi Toolkit, sedangkan Transformer Vanilla diimplementasikan menggunakan Python. Evaluasi dilakukan menggunakan Word Error Rate (WER) dan Character Error Rate (CER). Hasil penelitian menunjukkan bahwa Transformer Vanilla memiliki performa terbaik dengan WER sebesar 0,0222 dan CER sebesar 0,0164, dibandingkan HMM, HMM-GMM, dan HMM-GMM-DNN. Hasil ini menunjukkan bahwa Transformer Vanilla lebih efektif dalam mengenali ucapan bahasa Indonesia pada dataset penelitian.
Downloads
References
M. M. Reimann, F. A. Kunneman, C. Oertel, and K. V. Hindriks, “A Survey on Dialogue Management in Human-robot Interaction,” ACM Trans. Hum. Robot. Interact., vol. 13, no. 2, Jun. 2024, doi: 10.1145/3648605.
Z. Janeczko and M. E. Foster, “A Study on Human Interactions with Robots Based on Their Appearance and Behaviour,” in ACM International Conference Proceeding Series, Association for Computing Machinery, Jul. 2022. doi: 10.1145/3543829.3544523.
N. Mawaddah, C. R. Tamba, D. Hutagalung, F. Nainggolan, E. Lubis, and E. Jamzuri, “Automatic Speech Recognition for Human-Robot Interaction on The Humanoid Robot Barelang 7,” European Alliance for Innovation n.o., Feb. 2024. doi: 10.4108/eai.7-11-2023.2342940.
T. F. Abidin, A. Misbullah, R. Ferdhiana, L. Farsiah, M. Z. Aksana, and H. Riza, “Acoustic Model with Multiple Lexicon Types for Indonesian Speech Recognition,” Applied Computational Intelligence and Soft Computing, vol. 2022, 2022, doi: 10.1155/2022/3227828.
K. Nugroho, E. Noersasongko, Purwanto, Muljono, and D. R. I. M. Setiadi, “Enhanced Indonesian Ethnic Speaker Recognition using Data Augmentation Deep Neural Network,” Journal of King Saud University - Computer and Information Sciences, vol. 34, no. 7, pp. 4375–4384, Jul. 2022, doi: 10.1016/j.jksuci.2021.04.002.
A. Gulati et al., “Conformer: Convolution-augmented Transformer for Speech Recognition,” May 2020, [Online]. Available: http://2005.08100
A. Loubser, P. De Villiers, and A. De Freitas, “End-to-end automated speech recognition using a character based small scale transformer architecture,” Expert Syst. Appl., vol. 252, Oct. 2024, doi: 10.1016/j.eswa.2024.124119.
A. Adila, D. Lestari, A. Purwarianti, D. Tanaya, K. Azizah, and S. Sakti, “Enhancing Indonesian Automatic Speech Recognition: Evaluating Multilingual Models with Diverse Speech Variabilities,” Oct. 2024, [Online]: http://2410.08828
R. Atika, S. Dwijayanti, and B. Y. Suprapto, “Improving Speech-To-Text For The Indonesian Language Using A Modified Transformer,” Eastern-European Journal of Enterprise Technologies, vol. 1, no. 9 (139), pp. 78–90, 2026, doi: 10.15587/1729-4061.2026.350949.
R. Atika, S. Dwijayanti, and B. Y. Suprapto, “Comparing Transformer Models for Indonesian Speech-to-Text: Vanilla and Whisper,” in 2025 International Conference on Artificial Intelligence and Technological Solutions: For Good Health, Well-Being, and Sustainable Water Management in Support of SDGs 3, 6, and 9, ICAITech 2025 - Proceeding, Institute of Electrical and Electronics Engineers Inc., 2025, pp. 328–334. doi: 10.1109/ICAITech66481.2025.11387506.
H. Ahlawat, N. Aggarwal, and D. Gupta, “Automatic Speech Recognition: A survey of deep learning techniques and approaches,” Dec. 01, 2025, KeAi Communications Co. doi: 10.1016/j.ijcce.2024.12.007.
T. F. Abidin, A. Misbullah, R. Ferdhiana, L. Farsiah, M. Z. Aksana, and H. Riza, “Acoustic Model with Multiple Lexicon Types for Indonesian Speech Recognition,” Applied Computational Intelligence and Soft Computing, vol. 2022, 2022, doi: 10.1155/2022/3227828.
C. Park, M. Chen, and T. Hain, “Automatic Speech Recognition System-Independent Word Error Rate Estimation,” May 2024.
P. Chhabra and S. Goyal, “A Thorough Review on Deep Learning Neural Network,” Jan. 2023.
A. Vaswani et al., “Attention Is All You Need,” Aug. 2023, [Online]. Available: http://1706.03762
A. R. Sajun, I. Zualkernan, and D. Sankalpa, “A Historical Survey of Advances in Transformer Architectures,” May 01, 2024, Multidisciplinary Digital Publishing Institute (MDPI). doi: 10.3390/app14104316.
S. Zhang, J. Hale, M. Renwick, Z. Vrzić, and K. Langston, “An Evaluation of Croatian ASR Models for Čakavian Transcription,” May 2024.
J. Linke, S. Wepner, G. Kubin, and B. Schuppler, “Using Kaldi for Automatic Speech Recognition of Conversational Austrian German,” Jan. 2023, [Online]. Available: http://2301.06475
M. Müller, “Leveraging Python in AI and Machine Learning: A Survey of Techniques and Educational Approaches in Software Engineering,” Jan. 2025. [Online]. Available: www.ijeais.org/ijeais

This work is licensed under a Creative Commons Attribution-NonCommercial 4.0 International License.



