Implementation of the Indobert Model for Named Entity Recognition in Public Complaint Texts from the City of Semarang
Keywords:
IndoBERT, Named Entity Recognition, Public Complaints, Semarang City, Smart Governance.Abstract
Public complaint management in public services continues to face challenges because complaint data are generally presented as unstructured text containing informal language, abbreviations, and Indonesian-Javanese code-mixing. These characteristics complicate the identification of essential information and the determination of appropriate destination agencies when complaints are processed manually. This study aims to analyze the capability of the IndoBERT model for Named Entity Recognition (NER) in extracting relevant entities from public complaint texts in Semarang City to support complaint routing. A quantitative experimental approach was employed through complaint data collection, BIO-scheme annotation, text preprocessing, IndoBERT fine-tuning, and model evaluation using Precision, Recall, and F1-Score. The dataset consisted of 1,066 complaint sentences containing four target entities: service, location, problem category, and time. Strict-chunk evaluation using seqeval showed that the best IndoBERT checkpoint, obtained at Epoch 3, achieved a Macro F1-Score of 0.62, with the highest performance for LAYANAN (F1=0.68) and the lowest for LOKASI (F1=0.56). Error analysis further indicated that entity-boundary shifts, category ambiguity, and lexical variation in local and informal expressions remained important factors affecting recognition performance. The findings demonstrate that IndoBERT can transform heterogeneous complaint narratives into structured entity information and provide a foundation for improving the efficiency of public complaint management and routing within smart governance systems.
Downloads
References
Agustina, N., Naseer, M., Gusdevi, H., & Rismayadi, D. A. (2024, November). Development of a public complaint classification model to support E-Government using IndoBERT. In 2024 6th International Conference on Cybernetics and Intelligent System (ICORIS) (pp. 1-6). IEEE. https://doi.org/10.1109/ICORIS63540.2024.10903819
Aji, S., Maryani, I., & Muningsih, E. (2022). Analisis sentiment masyarakat menggunakan penggabungan algoritma Naive Bayes dan Particle Swarm Optimization. IJCIT (Indonesian Journal on Computer and Information Technology), 7(2), 119–126. https://doi.org/10.31294/ijcit.v7i2.14086
Alshammari, N., & Alanazi, S. (2021). The impact of using different annotation schemes on named entity recognition. Egyptian Informatics Journal, 22(3), 295–302. https://doi.org/10.1016/j.eij.2020.10.004
Asri, Y., Kuswardani, D., Suliyanti, W. N., & Fadhilah, N. A. (2025, October). A Domain-Specific Framework for Sentiment Analysis and Named Entity Recognition of PLN Mobile User Complaints. In 2025 International Conference on Advanced Technologies in Energy and Informatic (ICATEI) (pp. 597-602). IEEE. https://doi.org/10.1109/ICATEI67676.2025.11405393
Budi, I., & Suryono, R. R. (2023). Application of named entity recognition method for Indonesian datasets: A review. Bulletin of Electrical Engineering and Informatics, 12(2), 969–978. https://doi.org/10.11591/eei.v12i2.4529
Cahyawijaya, S., Winata, G. I., Wilie, B., Vincentio, K., Li, X., Kuncoro, A., Ruder, S., Lim, Z. Y., Bahar, S., Khodra, M., Purwarianti, A., & Fung, P. (2021). IndoNLG: Benchmark and resources for evaluating Indonesian natural language generation. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing (pp. 8875–8898). Association for Computational Linguistics. https://doi.org/10.18653/v1/2021.emnlp-main.699
Cahyo, P. W., Aesyi, U. S., Setianto, W. A., & Sulaiman, T. (2025). A novel named entity recognition approach of Indonesian fake news using part of speech and BERT model on presidential election. International Journal of Information Management Data Insights, 5(2), 100354. https://doi.org/10.1016/j.jjimei.2025.100354
Erna, E. D., Baskoro, D., & Firliana, R. (2026). Implementation of the IndoBERT Model on Named Entity Recognition for Entity Identification in the Folklore of the Tale of Wayang Arjuna & Purusara. The Indonesian Journal of Computer Science Research, 5(2), 131-137. https://doi.org/10.59095/ijcsr.v5i2.283
Firdaus, M. P., & Trisnawarman, D. (2025). Public sentiment analysis of the public housing savings program using the IndoBERT Lite model on YouTube comments. MALCOM: Indonesian Journal of Machine Learning and Computer Science, 5(1), 359–368. https://doi.org/10.57152/malcom.v5i1.1744
Grandini, M., Bagli, E., & Visani, G. (2020). Metrics for multi-class classification: An overview. arXiv. https://arxiv.org/abs/2008.05756
Hidayat, R., & Gata, W. (2025). Pemanfaatan IndoBERT untuk analisis sentimen ulasan aplikasi Depok Single Window. Information System for Educators and Professionals: Journal of Information System, 10(1), 13–24. https://doi.org/10.51211/isbi.v10i1.3334
Hidayatullah, A. F., Apong, R. A., Lai, D. T. C., & Qazi, A. (2023). Corpus creation and language identification for code-mixed Indonesian-Javanese-English tweets. PeerJ Computer Science, 9, e1312. https://doi.org/10.7717/peerj-cs.1312
Istiqomah, N., & Novika, F. (2025). Perbandingan kinerja model NER IndoBERT dan IndoLEM dalam ekstraksi informasi kesehatan pascabencana dari berita daring di Indonesia: Comparative performance of IndoBERT and IndoLEM baseline models for post-disaster health information extraction from Indonesian online news. Journal of Computer Science and Informatics Engineering, 4(3), 158–174. https://doi.org/10.55537/cosie.v4i3.1173
Maharani, W., & Latifa, A. Z. (2025). Analyzing public sentiment on the relocation of Indonesia's capital to Kalimantan as the Ibu Kota Nusantara using logistic regression. Jurnal Teknik Informatika (JUTIF), 6(2), 575–592. https://doi.org/10.52436/1.jutif.2025.6.2.4230
Pitaloka Subagyo, A. D., Nurcahyanto, H., & Marom, A. (2023). Analisis sistem pengelolaan pengaduan masyarakat pada Pusat Pengelolaan Pengaduan Masyarakat (P3M) Kota Semarang. Journal of Management and Public Policy, 12(2), 621–634. https://doi.org/10.14710/jppmr.v12i2.38488
Putra Pratama, I. W. B. S., Putra, M. A. P., & Paramitha, A. A. I. I. (2026). Deteksi berita hoaks bahasa Indonesia menggunakan model hybrid IndoBERT-BiLSTM dengan teknik hyperparameter tuning. JATI (Jurnal Mahasiswa Teknik Informatika), 10(2), 2830–2838. https://doi.org/10.36040/jati.v10i2.17872
Tandi, T. Y., Abidin, T. F., & Riza, H. (2025). Incorporation of IndoBERT and machine learning features to improve the performance of Indonesian textual entailment recognition. Journal of Information Systems Engineering and Business Intelligence, 11(2), 173–186. https://doi.org/10.20473/jisebi.11.2.173-186
Umam, A. K. ., Alzami, F., Sani, R. R. ., Rohmani, A. ., Prabowo, D. P. ., Pergiwati, D. ., Megantara, R. A. ., & Iswahyudi, I. (2025). Enhancing Entity Extraction in E-Government Complaint Data using LDA-Assisted NER. Sinkron : Jurnal Dan Penelitian Teknik Informatika, 9(4), 1878-1888. https://doi.org/10.33395/sinkron.v9i4.15292
Utomo, V. G. (2025). Benchmarking IndoBERT and Transformer Models for Sentiment Classification on Indonesian E-Government Service Reviews. Jurnal Transformatika, 23(1), 86-95. https://doi.org/10.26623/transformatika.v23i1.12095
Yulianti, E., Bhary, N., Abdurrohman, J., Dwitilas, F. W., Nuranti, E. Q., & Husin, H. S. (2024). Named entity recognition on Indonesian legal documents: A dataset and study using transformer-based models. International Journal of Electrical and Computer Engineering, 14(5), 5489–5501. https://doi.org/10.11591/ijece.v14i5.pp5489-5501









