FINE-TUNING MODEL BAHASA BESAR OPEN-SOURCE UNTUK LITERASI DAN KREATIVITAS INDONESIA: EVALUASI KOMPARATIF
Abstract
Adaptasi Model Bahasa Besar (LLM) ke dalam konteks budaya non-Inggris, terutama dalam bidang yang sarat nuansa seperti literasi dan kreativitas, masih menjadi tantangan yang signifikan. Penelitian ini mengatasi kesenjangan tersebut dengan mengembangkan dan mengevaluasi ‘Cakrakarsa’, sebuah chatbot AI yang dibuat melalui penyempurnaan model open-source Mistral-7B-Instruct. Model ini diadaptasi menggunakan teknik QLoRA pada dataset yang dikurasi khusus berisi 11.330 entri yang mencakup cerita rakyat, sastra, dan puisi Indonesia. Kinerja diukur terhadap model dasar melalui kerangka evaluasi internal yang komprehensif, yang menggabungkan metrik kuantitatif (Perplexity, Evaluation Loss) dan penilaian kualitatif melalui metode LLM-as-a-Judge pada 30 prompt khusus. Model yang telah disesuaikan menunjukkan keunggulan yang signifikan, dengan pengurangan sebesar 67,11% pada Perplexity dan 61,52% pada Evaluation Loss. Secara kualitatif, model inimenang dalam 81,11% perbandingan langsung, menunjukkan peningkatan yangmencolok dalam pemahaman konteks budaya, koherensi, dan kreativitas. Temuan kritis mengungkapkan adanyapertukaran antara spesialisasi domain yang mendalam dan pengetahuan umum yang luas. Penelitian inimenyumbangkan chatbot yang fungsional dan sadar budaya, serta metodologi yang kokoh dan dapat direplikasi untuk menyesuaikan LLM dengan domain budaya spesifik dengan sumber daya komputasi yang terbatas.
The adaptation of Large Language Models (LLMs) to non-English cultural contexts, particularly in nuanced domains like literacy and creativity, remains a significant challenge. This study addresses this gap by developing and evaluating 'Cakrakarsa', an AI chatbot created by fine-tuning the open-source Mistral-7B-Instruct model. The model was adapted using the QLoRA technique on a custom-curated dataset of 11,330 entries comprising Indonesian folklore, literature, and poetry. Performance was measured against the baseline model through a comprehensive internal evaluation framework, combining quantitative metrics (Perplexity, Evaluation Loss) and qualitative assessment via the LLM-as-a-Judge method on 30 dedicated prompts. The fine-tuned model demonstrated significant superiority, achieving a 67.11% reduction in Perplexity and a 61.52% reduction in Evaluation Loss. Qualitatively, it won 81.11% of head-to-head comparisons, showing marked improvements in cultural context understanding, coherence, and creativity. A critical finding reveals a trade-off between deep domain specialization and broad general knowledge. This research contributes a functional, culturally-aware chatbot and a robust, replicable methodology for adapting LLMs to specific cultural domains with limited computational resources.
Downloads
References
H. Naveed et al., “A Comprehensive Overview of Large Language Models,” ACM Trans. Intell. Syst. Technol., vol. 16, no. 5, Aug. 2025, doi: 10.1145/3744746.
T. B. Brown et al., “Language Models are Few-Shot Learners,” in Proceedings of the 34th International Conference on Neural Information Processing Systems, Red Hook, NY, USA: Curran Associates, Inc., 2020, pp. 1877–1901. [Online]. Available: https://proceedings.neurips.cc/paper_files/paper/2020/file/1457c0d6bfcb4967418bfb8ac142f64a-Paper.pdf
Y. Tao, O. Viberg, R. S. Baker, and R. F. Kizilcec, “Cultural Bias and Cultural Alignment of Large Language Models,” Jun. 2024, doi: 10.1093/pnasnexus/pgae346.
WIPO, “Global Innovation Index 2023,” 2023.
OECD, “PISA 2022 Results Factsheets Indonesia PUBE,” 2023. [Online]. Available: https://oecdch.art/a40de1dbaf/C108.
J. Wei et al., “Finetuned Language Models Are Zero-Shot Learners,” in International Conference on Learning Representations (2022), 2022. [Online]. Available: https://openreview.net/forum?id=gEZrGCozdqR
C. Jeong, “Domain-specialized LLM: Financial fine-tuning and utilization method using Mistral 7B,” Journal of Intelligence and Information Systems, vol. 30, pp. 93–120, Sep. 2024, doi: 10.13088/jiis.2024.30.1.093.
Y. Zhao et al., “Assessing and Understanding Creativity in Large Language Models,” Machine Intelligence Research, vol. 22, no. 3, pp. 417–436, 2025, doi: 10.1007/s11633-025-1546-4.
E. J. Hu et al., “LoRA: Low-Rank Adaptation of Large Language Models,” in International Conference on Learning Representations, 2022. Accessed: Sep. 07, 2025. [Online]. Available: https://openreview.net/forum?id=nZeVKeeFYf9
T. Dettmers, A. Pagnoni, A. Holtzman, and L. Zettlemoyer, “QLoRA: Efficient Finetuning of Quantized LLMs,” in Advances in Neural Information Processing Systems, A. Oh, T. Naumann, A. Globerson, K. Saenko, M. Hardt, and S. Levine, Eds., Curran Associates, Inc., 2023, pp. 10088–10115. [Online]. Available: https://proceedings.neurips.cc/paper_files/paper/2023/file/1feb87871436031bdc0f2beaa62a049b-Paper-Conference.pdf
S. Liu et al., “Judge as A Judge: Improving the Evaluation of Retrieval-Augmented Generation through the Judge-Consistency of Large Language Models,” in Findings of the Association for Computational Linguistics: ACL 2025, W. Che, J. Nabende, E. Shutova, and M. T. Pilehvar, Eds., Vienna, Austria: Association for Computational Linguistics, Jul. 2025, pp. 5788–5807. doi: 10.18653/v1/2025.findings-acl.301.
G. H. Chen, S. Chen, Z. Liu, F. Jiang, and B. Wang, “Humans or LLMs as the Judge? A Study on Judgement Bias,” in Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, Y. Al-Onaizan, M. Bansal, and Y.-N. Chen, Eds., Miami, Florida, USA: Association for Computational Linguistics, Nov. 2024, pp. 8301–8327. doi: 10.18653/v1/2024.emnlp-main.474.
R. Y. Pratama and Supatman, “Pengembangan Sistem Chatbot Cerdas Berbasis Natural Language Processing (NLP) untuk Peningkatan Layanan Informasi Hotel,” JITET (Jurnal Informatika dan Teknik Elektro Terapan), vol. 14, no. 1, pp. 1350–1362, 2025, doi: 10.23960/jitet.v14i1.8528.
W.-L. Chiang et al., “Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference,” in Proceedings of the 41st International Conference on Machine Learning, R. Salakhutdinov, Z. Kolter, K. Heller, A. Weller, N. Oliver, J. Scarlett, and F. Berkenkamp, Eds., Vienna, Austria: JMLR.org, 2024, pp. 8359–8388. [Online]. Available: https://proceedings.mlr.press/v235/chiang24b.html
A. Malik, S. Mayhew, C. Piech, and K. Bicknell, “From Tarzan to Tolkien: Controlling the Language Proficiency Level of LLMs for Content Generation,” in Findings of the Association for Computational Linguistics: ACL 2024, L.-W. Ku, A. Martins, and V. Srikumar, Eds., Bangkok, Thailand: Association for Computational Linguistics, Aug. 2024, pp. 15670–15693. doi: 10.18653/v1/2024.findings-acl.926.
Daniel Han, Michael Han, and Team, “Unsloth,” 2023. Accessed: Feb. 14, 2024. [Online]. Available: https://github.com/unslothai/unsloth

This work is licensed under a Creative Commons Attribution-NonCommercial 4.0 International License.



