Deteksi Hoaks Berita Berbahasa Indonesia Menggunakan IndoBERT dengan Penanganan Ketidakseimbangan Kelas Berbasis Gabungan Class Weighting dan Focal Loss
DOI:
https://doi.org/10.29408/jit.v9i2.35214Keywords:
class imbalance, focal loss, hoax detection, IndoBERT, natural language processingAbstract
The spread of hoaxes in Indonesian-language digital media has risen sharply in the past five years with wide societal harm. Most prior work on Indonesian hoax detection emphasizes model architecture, while the class-imbalance problem in field data receives less attention. This study presents a data-centric approach combining the IndoBERT pretrained language model with a hybrid class-imbalance objective of class weighting and focal loss. The dataset is the indonesiafalsenews corpus of 4231 labelled articles with an approximate 4.5 to 1 hoax-to-fact ratio. Evaluation uses five-seed runs with McNemar and paired-bootstrap significance testing. The proposed configuration achieves an F1-macro of 0.7240, accuracy of 0.8392, ROC-AUC of 0.8083, and a Matthews correlation coefficient of 0.4510, significantly outperforming every TF-IDF baseline at the 0.05 level. Text augmentation does not provide consistent gains, which implies that imbalance handling is the more effective lever for Indonesian hoax detection.
References
1] M. F. Mridha, A. J. Keya, Md. A. Hamid, M. M. Monowar, and Md. S. Rahman, “A Comprehensive Review on Fake News Detection With Deep Learning,” IEEE Access, vol. 9, pp. 156151–156170, 2021, doi: 10.1109/ACCESS.2021.3129329.
[2] F. Koto, A. Rahimi, J. H. Lau, and T. Baldwin, “IndoLEM and IndoBERT: A Benchmark Dataset and Pre-trained Language Model for Indonesian NLP,” in Proceedings of the 28th International Conference on Computational Linguistics, Barcelona, Spain (Online): International Committee on Computational Linguistics, 2020, pp. 757–770, doi: 10.18653/v1/2020.coling-main.66.
[3] B. Wilie et al., “IndoNLU: Benchmark and Resources for Evaluating Indonesian Natural Language Understanding,” in Proceedings of the 1st Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics and the 10th International Joint Conference on Natural Language Processing, Suzhou, China: Association for Computational Linguistics, 2020, pp. 843–857, doi: 10.18653/v1/2020.aacl-main.85.
[4] M. Y. Ridho, and E. Yulianti, “From Text to Truth: Leveraging IndoBERT and Machine Learning Models for Hoax Detection in Indonesian News,” Jurnal Ilmiah Teknik Elektro Komputer dan Informatika (JITEKI), vol. 10, no. 3, pp. 544–555, 2024.
[5] D. Y. Yefferson, V. Lawijaya, and A. S. Girsang, “Hybrid model: IndoBERT and long short-term memory for detecting Indonesian hoax news,” IAES International Journal of Artificial Intelligence (IJ-AI), vol. 13, no. 2, pp. 1913–1924, 2024, doi: 10.11591/ijai.v13.i2.pp1913-1924.
[6] S. Henning, W. Beluch, A. Fraser, and A. Friedrich, “A Survey of Methods for Addressing Class Imbalance in Deep-Learning Based Natural Language Processing,” 2023, pp. 523–540, doi: 10.18653/v1/2023.eacl-main.38.
[7] M. Bayer, M.-A. Kaufhold, and C. Reuter, “A Survey on Data Augmentation for Text Classification,” ACM Computing Surveys, vol. 55, no. 7, p. 146, 2022, doi: 10.1145/3544558.
[8] D. Zha et al., “Data-centric Artificial Intelligence: A Survey,” ACM Computing Surveys, vol. 57, no. 5, 2025, doi: 10.1145/3711118.
[9] M. I. K. Sinapoy, Y. Sibaroni, and S. S. Prasetyowati, “Comparison of LSTM and IndoBERT Method in Identifying Hoax on Twitter,” Jurnal RESTI (Rekayasa Sistem dan Teknologi Informasi), vol. 7, no. 3, pp. 657–662, 2023.
[10] F. Muftie, and M. Haris, “IndoBERT Based Data Augmentation for Indonesian Text Classification,” 2023, pp. 128–132, https://ieeexplore.ieee.org/document/10250061/.
[11] J. Tian et al., “Re-embedding Difficult Samples via Mutual Information Constrained Semantically Oversampling for Imbalanced Text Classification,” 2021, doi: 10.18653/v1/2021.emnlp-main.252.
[12] K. R. M. Fernando, and C. P. Tsokos, “Dynamically Weighted Balanced Loss: Class Imbalanced Learning and Confidence Calibration of Deep Neural Networks,” IEEE Transactions on Neural Networks and Learning Systems, vol. 33, no. 7, pp. 2940–2951, 2022, doi: 10.1109/TNNLS.2020.3047335.
[13] T.-Y. Lin, P. Goyal, R. Girshick, K. He, and P. Dollár, “Focal Loss for Dense Object Detection,” in Proceedings of the IEEE International Conference on Computer Vision (ICCV 2017), 2017, pp. 2980–2988, doi: 10.1109/ICCV.2017.324.
[14] Y. Cui, M. Jia, T.-Y. Lin, Y. Song, and S. Belongie, “Class-Balanced Loss Based on Effective Number of Samples,” 2019, pp. 9268–9277, doi: 10.1109/CVPR.2019.00949.
[15] K. Cao, C. Wei, A. Gaidon, N. Arechiga, and T. Ma, “Learning Imbalanced Datasets with Label-Distribution-Aware Margin Loss,” 2019, pp. 1567–1578, https://proceedings.neurips.cc/paper/2019/hash/621461af90cadfdaf0e8d4cc25129f91-Abstract.html.
[16] J. Chen, D. Tam, C. Raffel, M. Bansal, and D. Yang, “An Empirical Survey of Data Augmentation for Limited Data Learning in NLP,” Transactions of the Association for Computational Linguistics, vol. 11, pp. 191–211, 2023, doi: 10.1162/tacl_a_00542.
[17] T. G. Dietterich, “Approximate Statistical Tests for Comparing Supervised Classification Learning Algorithms,” Neural Computation, vol. 10, no. 7, pp. 1895–1923, 1998, doi: 10.1162/089976698300017197.
[18] D. Ulmer, C. Hardmeier, and J. Frellsen, “deep-significance: Easy and Meaningful Statistical Significance Testing in the Age of Neural Networks,” 2022, doi: 10.48550/arXiv.2204.06815.
[19] J. Dodge, S. Gururangan, D. Card, R. Schwartz, and N. A. Smith, “Show Your Work: Improved Reporting of Experimental Results,” 2019, pp. 2185–2194, doi: 10.18653/v1/D19-1224.
[20] D. Chicco, and G. Jurman, “The advantages of the Matthews correlation coefficient (MCC) over F1 score and accuracy in binary classification evaluation,” BMC Genomics, vol. 21, p. 6, 2020, doi: 10.1186/s12864-019-6413-7.
Downloads
Published
How to Cite
Issue
Section
License
Copyright (c) 2026 Infotek: Jurnal Informatika dan Teknologi

This work is licensed under a Creative Commons Attribution 4.0 International License.
Semua tulisan pada jurnal ini menjadi tanggung jawab penuh penulis. Jurnal Infotek memberikan akses terbuka terhadap siapapun agar informasi dan temuan pada artikel tersebut bermanfaat bagi semua orang. Jurnal Infotek ini dapat diakses dan diunduh secara gratis, tanpa dipungut biaya sesuai dengan lisense creative commons yang digunakan.
Jurnal Infotek is licensed under a Creative Commons Attribution 4.0 International License.
Statistik Pengunjung


