Perbandingan Akurasi Model AI Dalam Menilai Kepribadian Pengguna Menggunakan Pendekatan NLP (Natural Language Processing)

  • Ryan Hadi Ardhiansyah
    Universitas Esa Unggul
DOI: https://doi.org/10.23960/jitet.v14i3.9539
Keywords Kecerdasan Buatan, Pemrosesan Bahasa Alami, Lima Dimensi Kepribadian Utama, Pembelajaran Mendalam, Model Berbasis Aturan
Abstract Views (Last 12 Months)
6 Abstract Views
4 Downloads

Abstract

Penilaian kepribadian berbasis teks merupakan salah satu aplikasi NLP yang berkembang pesat, namun pemilihan model yang tepat antara efisiensi dan akurasi masih menjadi tantangan. Penelitian ini membandingkan efektivitas pendekatan deep learning (BERT dan DistilBERT) dengan metode rule-based dalam memprediksi lima dimensi kepribadian Big Five dari data teks. Kedua model transformer dioptimasi menggunakan Task-Specific Head Tuning dikombinasikan dengan kalibrasi threshold, dan dievaluasi pada dua kategori panjang teks: teks pendek (≤128 token) dan teks panjang (≥256 token) dari dataset agregat sebanyak 37.630 sampel. Hasil menunjukkan bahwa DistilBERT dengan full fine-tuning mencapai performa terbaik pada teks pendek (F1-Score: 0,6751; nMCC: 0,6807), sementara BERT dengan full fine-tuning unggul pada teks panjang (nMCC: 0,6258). Metode rule-based secara konsisten menghasilkan performa lebih rendah dengan nMCC tertinggi hanya 0,5288. Analisis per dimensi menunjukkan bahwa Openness paling mudah dideteksi pada teks pendek (F1-Score: 0,8405), sedangkan Agreeableness membutuhkan konteks lebih panjang (F1-Score: 0,7198). Penelitian ini menyimpulkan bahwa DistilBERT merupakan pilihan optimal untuk lingkungan dengan sumber daya komputasi terbatas, sementara BERT tetap unggul untuk analisis teks panjang yang kompleks.

 

Text-based personality assessment is a rapidly growing NLP application, yet selecting the appropriate model that balances efficiency and accuracy remains challenging. This study compares the effectiveness of deep learning approaches (BERT and DistilBERT) against a rule-based method in predicting Big Five personality traits from text data. Both transformer models were optimized using Task-Specific Head Tuning combined with threshold calibration, and evaluated on two text length categories: short texts (≤128 tokens) and long texts (≥256 tokens) from an aggregated dataset of 37,630 samples. Results indicate that DistilBERT with full fine-tuning achieved the best performance on short texts (F1-Score: 0.6751; nMCC: 0.6807), while BERT with full fine-tuning outperformed on long texts (nMCC: 0.6258). Rule-based methods consistently underperformed with the highest nMCC reaching only 0.5288. Per-dimension analysis revealed that Openness was most easily detected on short texts (F1-Score: 0.8405), whereas Agreeableness required longer context (F1-Score: 0.7198). This study concludes that DistilBERT is the optimal choice for resource-constrained environments, while BERT remains superior for analyzing long and contextually complex texts.

Downloads

Download data is not yet available.

References

F. Celli dan B. Lepri, “Is Big Five better than MBTI?,” dalam Proceedings of the Fifth Italian Conference on Computational Linguistics CLiC-it 2018, Accademia University Press, 2019, hlm. 93–98. doi: 10.4000/books.aaccademia.3147.

Y. Mehta, N. Majumder, A. Gelbukh, dan E. Cambria, “Recent Trends in Deep Learning Based Personality Detection,” Agu 2019, doi: 10.1007/s10462-019-09770-z.

G. Serapio-García dkk., “Personality Traits in Large Language Models,” Jun 2023, doi: 10.48550/arXiv.2307.00184.

J. Devlin, M.-W. Chang, K. Lee, K. T. Google, dan A. I. Language, “BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding,” 2019. doi: 10.18653/v1/N19-1423.

H. Christian, D. Suhartono, A. Chowanda, dan K. Z. Zamli, “Text based personality prediction from multiple social media data sources using pre-trained language model and model averaging,” J. Big Data, vol. 8, no. 1, Des 2021, doi: 10.1186/s40537-021-00459-1.

V. Sanh, L. Debut, J. Chaumond, dan T. Wolf, “DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter,” Mar 2020, doi: 10.48550/arXiv.1910.01108.

N. Fatima, S. Gul, J. Ahmed, Z. H. Khand, dan G. Mujtaba, “A rule-based machine learning model for career selection through MBTI personality,” Mehran University Research Journal of Engineering and Technology, vol. 41, no. 2, hlm. 185–196, Apr 2022, doi: 10.22581/muet1982.2202.18.

N. Koenig dkk., “Improving measurement and prediction in personnel selection through the application of machine learning,” Pers. Psychol., vol. 76, no. 4, hlm. 1061–1123, Des 2023, doi: 10.1111/peps.12608.

E. Kerz, Y. Qiao, S. Zanwar, dan D. Wiechmann, “Pushing on Personality Detection from Verbal Behavior: A Transformer Meets Text Contours of Psycholinguistic Features,” 2022. doi: 10.18653/v1/2022.wassa-1.17.

Dr. V Geetha, D. C. K. Gomathy, Mr. D. S. D. Vallab Yaratha Yagn, dan S. Praneesh, “THE ROLE OF NATURAL LANGUAGE PROCESSING,” INTERANTIONAL JOURNAL OF SCIENTIFIC RESEARCH IN ENGINEERING AND MANAGEMENT, vol. 07, no. 11, hlm. 1–11, Nov 2023, doi: 10.55041/IJSREM27094.

A. Bittermann dan A. Fischer, “Natural Language Processing in Psychology,” 1 Juli 2024, Hogrefe Verlag GmbH & Co. KG. doi: 10.1027/2151-2604/a000568.

Y. Lecun, Y. Bengio, dan G. Hinton, “Deep learning,” 27 Mei 2015, Nature Publishing Group. doi: 10.1038/nature14539.

R. R. Mccrae dan O. P. John, “An Introduction to the Five-Factor Model and Its Applications,” 1992. doi: 10.1111/j.1467-6494.1992.tb00970.x.

M. Sokolova dan G. Lapalme, “A systematic analysis of performance measures for classification tasks,” Inf. Process. Manag., vol. 45, no. 4, hlm. 427–437, Jul 2009, doi: 10.1016/j.ipm.2009.03.002.

D. Chicco dan G. Jurman, “The advantages of the Matthews correlation coefficient (MCC) over F1 score and accuracy in binary classification evaluation,” BMC Genomics, vol. 21, no. 1, Jan 2020, doi: 10.1186/s12864-019-6413-7.

M. Gjurković, M. Karan, I. Vukojević, M. Bošnjak, dan J. Šnajder, “PANDORA Talks: Personality and Demographics on Reddit,” Jun 2021, doi: 10.18653/v1/2021.socialnlp-1.12.

F. Celli, F. Pianesi, D. Stillwell, dan M. Kosinski, “Workshop on Computational Personality Recognition: Shared Task,” 2021. doi: 10.1609/icwsm.v7i2.14467.

J. W. Pennebaker dan L. A. King, “Linguistic Styles: Language Use as an Individual Difference,” 1999. doi: 10.1037/0022-3514.77.6.1296.

H. Pouransari dkk., “Dataset Decomposition: Faster LLM Training with Variable Sequence Length Curriculum,” Jan 2025, doi: 10.48550/arXiv.2405.13226.

L. Xu, H. Xie, S.-Z. J. Qin, X. Tao, dan F. L. Wang, “Parameter-Efficient Fine-Tuning Methods for Pretrained Language Models: A Critical Review and Assessment,” Des 2023, doi: 10.48550/arXiv.2312.12148.

Z. Yang, M. Ding, Y. Guo, Q. Lv, dan J. Tang, “Parameter-Efficient Tuning Makes a Good Classification Head,” Mar 2023, doi: 10.48550/arXiv.2210.16771.

J. M. Digman, “PERSONALITY STRUCTURE: EMERGENCE OF THE FIVE-FACTOR MODEL,” 1990. doi: 10.1146/annurev.ps.41.020190.002221.

S. C. Ma, T. Ermakova, dan B. Fabian, “FairGridSearch: A Framework to Compare Fairness-Enhancing Models,” dalam Proceedings - 2023 22nd IEEE/WIC International Conference on Web Intelligence and Intelligent Agent Technology, WI-IAT 2023, Institute of Electrical and Electronics Engineers Inc., 2023, hlm. 388–394. doi: 10.1109/WI-IAT59888.2023.00064.

A. Chowanda, D. Suhartono, E. W. Andangsari, dan K. Z. bin Zamli, “MACHINE LEARNING ALGORITHMS EXPLORATION FOR PREDICTING PERSONALITY FROM TEXT,” ICIC Express Letters, vol. 16, no. 2, hlm. 117–125, Feb 2022, doi: 10.24507/icicel.16.02.117.

Divaretta Kiesa Sumartha, “IMPLEMENTASI INDOBERT UNTUK ANALISIS SENTIMEN OPINI PUBLIK TERHADAP KEBIJAKAN KENAIKAN UKT DI ERA PEMERINTAHAN,” Jurnal Informatika dan Teknik Elektro Terapan, vol. 13, no. 3S1, Okt 2025, doi: 10.23960/jitet.v13i3S1.7880.

K. Kelvin dan Y. Utomo, “Overview of Text Based Personality Prediction Using Deep Learning,” Engineering, MAthematics and Computer Science Journal (EMACS), vol. 6, no. 2, hlm. 93–100, Mei 2024, doi: 10.21512/emacsjournal.v6i2.11550.

Cover
Published
2026-08-13
How to Cite
Ryan Hadi Ardhiansyah. (2026). Perbandingan Akurasi Model AI Dalam Menilai Kepribadian Pengguna Menggunakan Pendekatan NLP (Natural Language Processing). Jurnal Informatika Dan Teknik Elektro Terapan, 14(3). https://doi.org/10.23960/jitet.v14i3.9539