Comparing the Interpretability of Classical Machine Learning Models and IndoBERT in Sentiment Analysis of Indonesian Health Service Applications using Multi-Review Generalization Testing

Muhammad Ridlo Nu'man Hakim, Dwi Hosanna Bangkalang

Abstract


The increasing use of healthcare service applications in Indonesia has led to a growing volume of user reviews on the Google Play Store. This development makes manual sentiment analysis increasingly inefficient, particularly for reviews containing diverse linguistic expressions and contrastive sentences that involve sentiment shifts across clauses and are prone to causing classification errors. However, the ability of sentiment analysis models to handle such contrastive sentences has received limited attention. Moreover, previous sentiment analysis studies have tended to report performance metrics without systematically evaluating model generalization and interpretability. This study evaluates classical machine learning models, namely Naïve Bayes, Logistic Regression, Support Vector Machine (SVM), Random Forest, and XGBoost, alongside the IndoBERT transformer model for sentiment classification of Indonesian-language reviews of healthcare service applications, with particular emphasis on contrastive sentences. The evaluation covers performance comparison, out-of-distribution generalization testing, and interpretability analysis using Integrated Gradients. The research methodology includes collecting 20,000 reviews from the Google Play Store, labeling the reviews with validation against manually assigned labels, developing models using classical machine learning algorithms and fine-tuning IndoBERT, and conducting out-of-distribution generalization testing. The results show that IndoBERT achieved the highest F1-score of 98.10%, outperforming the classical machine learning models, which achieved F1-scores ranging from 96.28% to 97.57%. On contrastive reviews, IndoBERT demonstrated more consistent generalization performance, achieving an accuracy of 92.50% in the in-domain scenario and 77.50% in the cross-domain scenario, compared with ranges of 65.00%–72.50% and 67.50%–75.00%, respectively, for the classical models. Further analysis using Integrated Gradients revealed that IndoBERT's predictions were influenced more strongly by sentiment-bearing tokens, particularly adjectives and negations, than by contrastive conjunctions themselves. These findings provide a reference for developing Indonesian-language sentiment analysis models and support healthcare service application developers in evaluating service quality based on user perceptions.

Keywords


generalization test; IndoBERT; integrated gradients; interpretability; sentiment analysis

Full Text:

PDF

References


P. Julia and H. Ikhsan, “Transformasi Digital dalam Sistem Informasi Kesehatan: Dampak terhadap Kualitas Pelayanan medis,” Public Health and Safety International Journal, Vol. 5, No. 2, pp. 642–653, Oct. 2025, DOI: 10.55642/phasij.v5i02.1245.

B. Primin and A. P. Wibowo, “Implementasi Aplikasi berbasis Mobile untuk Pelayanan Jasa Kesehatan,” Jurnal Informatika: Jurnal Pengembangan IT, Vol. 8, No. 2, pp. 119–125, May 2023, DOI: 10.30591/jpit.v8i2.5076.

T. Sugihartono and R. R. C. Putra, “Penerapan Metode Support Vector Machine dalam Classifikasi Ulasan Pengguna Aplikasi Mobile JKN,” SKANIKA: Sistem Komputer dan Teknik Informatika, Vol. 7, No. 2, pp. 144–153, Jul. 2024, DOI: 10.36080/skanika.v7i2.3193.

S. Setianingsih, H. Hendri, B. O. Lubis, A. Sudradjat, I. Carolina, and W. Widiati, “Evaluasi Aplikasi SATUSEHAT dengan Metode Use Questionnaire dan IPA,” JATI (Jurnal Mahasiswa Teknik Informatika), Vol. 8, No. 3, pp. 2988–2995, May 2024, DOI: 10.36040/jati.v8i3.9416.

T. G. W. M. Sidabutar and D. Juardi, “Analisis Sentimen Masyarakat terhadap Penggunaan Halodoc sebagai Layanan Telemedicine di Indonesia,” Jurnal Informatika dan Teknik Elektro Terapan, Vol. 13, No. 1, pp. 571–577, Jan. 2025, DOI: 10.23960/jitet.v13i1.5682.

A. Putri, A. Faroqi, and S. Mukaromah, “Pendekatan Model UTAUT2 dalam Menilai Penerimaan Pengguna terhadap Layanan Telemedicine Alodokter,” KONSTELASI: Konvergensi Teknologi dan Sistem Informasi, Vol. 5, No. 1, pp. 141–154, Jun. 2025, DOI: 10.24002/konstelasi.v5i1.11696.

A. S. Berliana and M. Mustikasari, “Analisis Sentimen pada Ulasan Aplikasi JAKARTANOTEBOOK di Google Play menggunakan Metode Recurrent Neural Network (RNN),” Jurnal Informatika dan Teknik Elektro Terapan, Vol. 12, No. 3, pp. 3051–3057, Aug. 2024, DOI: 10.23960/jitet.v12i3.5067.

A. Komarudin and A. M. Hilda, “Analisis Sentimen Ulasan Aplikasi Identitas Kependudukan Digital pada Play Store menggunakan Metode Naïve Bayes,” Computer Science (CO-SCIENCE), Vol. 4, No. 1, pp. 28–36, Jan. 2024, DOI: 10.31294/coscience.v4i1.2955.

A. D. Fitriyanto and P. Purwanto, “Analisis Sentimen Ulasan DANA dari Play Store dengan Metode SVM, Logistic Regression, Naive Bayes dan KNN,” Building of Informatics, Technology and Science (BITS), Vol. 7, No. 3, pp. 1887–1899, Dec. 2025, DOI: 10.47065/bits.v7i3.8769.

F. M. Putra, P. W. Hardjita, and D. A. Tyas, “Sentimen Analisis Pengguna Media Sosial berdasarkan Metode Ekstraksi Fitur dan Klasifikasi,” Jurnal Ilmu Komputer, Vol. 16, No. 2, pp. 88–96, Sep. 2023, DOI: 10.24843/JIK.2023.v16.i02.p02.

E. Hokijuliandy, H. Napitupulu, and F. Firdaniza, “Analisis Sentimen menggunakan Metode Klasifikasi Support Vector Machine (SVM) dan Seleksi Fitur Chi-Square,” SisInfo : Jurnal Sistem Informasi dan Informatika, Vol. 5, No. 2, pp. 40–49, Aug. 2023, DOI: 10.37278/sisinfo.v5i2.670.

B. Mahendra, Martanto, D. Pratama, A. Faqih, and R. Kurniawan, “Evaluasi Pengaruh Kualitas Data terhadap Performa Model Machine Learning menggunakan Pendekatan Data-Centric AI,” Jurnal Sistem Informasi dan Teknologi (SINTEK), Vol. 6, No. 1, pp. 107–113, Jan. 2026, DOI: 10.56995/sintek.v6i1.211.

S. Biswas, K. Young, and J. Griffith, “A Comparison of Automatic Labelling Approaches for Sentiment Analysis,” in Proceedings of the 11th International Conference on Data Science, Technology and Applications, SCITEPRESS - Science and Technology Publications, Jul. 2022, pp. 312–319, DOI: 10.5220/0011265900003269.

P. Maxmiliano, Y. F. Riti, I. Y. Nugroho, and C. E. C. Juniarto, “Perbandingan Algoritma Naive Bayes dan Bert untuk Analisis Sentimen Ulasan Produk Shopee berdasarkan Rating dan Atribut Produk (Warna/Kategori),” Jurnal Media Informatika, Vol. 6, No. 5, pp. 2552–2565, Sep. 2025, DOI: 10.55338/jumin.v6i5.6761.

L. Kuswati and B. A. Habsy, “Analisis Sentimen Multi-Kelas Ulasan Aplikasi Satusehat menggunakan Indobert berbasis Transformer,” Jurnal Sosial dan Teknologi (SOSTECH), Vol. 6, No. 4, pp. 1540–1551, Apr. 2026, DOI: 10.59188/jurnalsostech.v6i4.32787.

H. Abriananta, K. Umam, N. C. H. Wibowo, and M. R. Handayani, “Comparison of Naive Bayes, Support Vector Machine, and Indobert Methods for Classifying Public Sentiment towards the MBG Program on Platform X,” Journal of Applied Informatics and Computing, Vol. 10, No. 3, pp. 2806–2815, Jun. 2026, DOI: 10.30871/jaic.v10i3.12721.

Paulia, C. Hualangi, and N. Sugiroto, “Kontrastif Konjungsi Transisi namun dan tetapi dalam Bahasa Indonesia dan Que dan Bu Guo dalam Bahasa Mandarin,” LINGUISTIK : Jurnal Bahasa & Sastra, Vol. 10, No. 2, pp. 171–185, Jun. 2025, DOI: 10.31604/linguistik.v10ii2.171-185.

H. Firda, P. Putra, N. R. Oktadini, P. E. Sevtiyuni, and A. Meiriza, “Comparison of Rating-based and Inset Lexicon-based Labeling in Sentiment Analysis using SVM (Case Study: GoBiz Application Reviews on Google Play Store),” SISTEMASI, Vol. 14, No. 2, pp. 516–529, Mar. 2025, DOI: 10.32520/stmsi.v14i2.4795.

A. S. Rizkia, W. Wufron, and F. F. Roji, “Analisis Sentimen Coretax: Perbandingan Pelabelan Data Manual, Transformers-based, dan Lexicon-based pada Performa IndoBERT,” MALCOM: Indonesian Journal of Machine Learning and Computer Science, Vol. 5, No. 3, pp. 1037–1048, Jul. 2025, DOI: 10.57152/malcom.v5i3.2151.

D. K. Tjong, S. Anwar, A. Hualangi, T. Agustina, and Anita, “Kontrastif Kata Penghubung Pertentangan dalam Bahasa Indonesia dan Bahasa Mandarin,” LINGUISTIK : Jurnal Bahasa dan Sastra, Vol. 11, No. 2, pp. 414–431, Jun. 2026, DOI: 10.31604/linguistik.v11i2.414-431.

Z. Chen, H. Sun, H. He, and P. Chen, “Learning from Noisy Crowd Labels with Logics,” in 2023 IEEE 39th International Conference on Data Engineering (ICDE), IEEE, Apr. 2023, pp. 41–52, DOI: 10.1109/ICDE55515.2023.00011.

D. Hupkes et al., “A Taxonomy and Review of Generalization Research in NLP,” Nat. Mach. Intell., Vol. 5, No. 10, pp. 1161–1174, Oct. 2023, DOI: 10.1038/s42256-023-00729-y.

H. Moraliyage, G. Kulawardana, D. De Silva, Z. Issadeen, M. Manic, and S. Katsura, “Explainable Artificial Intelligence with Integrated Gradients for the Detection of Adversarial Attacks on Text Classifiers,” Applied System Innovation, Vol. 8, No. 1, Feb. 2025, DOI: 10.3390/asi8010017.

N. K. Singh, K. Ghosh, J. Mahapatra, U. Garain, and A. Senapati, “HCDIR: End-to-end Hate Context Detection, and Intensity Reduction model for online comments,” Dec. 2023. DOI: 10.48550/arXiv.2312.13193.

M. Sundararajan, A. Taly, and Q. Yan, “Axiomatic Attribution for Deep Networks,” in Proceedings of the 34th International Conference on Machine Learning, Proceedings of Machine Learning Research, 2017, pp. 3319–3328.




DOI: https://doi.org/10.32520/stmsi.v15i9.6899

Article Metrics

Abstract view : 0 times
PDF - 0 times

Refbacks

  • There are currently no refbacks.


Creative Commons License
This work is licensed under a Creative Commons Attribution-ShareAlike 4.0 International License.