Deteksi Hoaks Berita Berbahasa Indonesia Menggunakan IndoBERT dengan Penanganan Ketidakseimbangan Kelas Berbasis Gabungan Class Weighting dan Focal Loss

Authors

  • Riadhul Muttaqin Universitas Islam Kalimantan Muhammad Arsyad Al Banjari Banjarmasin
  • Muhammad Edya Rosadi Universitas Islam Kalimantan Muhammad Arsyad Al Banjari Banjarmasin
  • Muhammad Iqbal Firdaus Universitas Islam Kalimantan Muhammad Arsyad Al Banjari Banjarmasin
  • Dian Agustini Universitas Islam Kalimantan Muhammad Arsyad Al Banjari Banjarmasin

DOI:

https://doi.org/10.29408/jit.v9i2.35214

Keywords:

class imbalance, focal loss, hoax detection, IndoBERT, natural language processing

Abstract

The spread of hoaxes in Indonesian-language digital media has risen sharply in the past five years with wide societal harm. Most prior work on Indonesian hoax detection emphasizes model architecture, while the class-imbalance problem in field data receives less attention. This study presents a data-centric approach combining the IndoBERT pretrained language model with a hybrid class-imbalance objective of class weighting and focal loss. The dataset is the indonesiafalsenews corpus of 4231 labelled articles with an approximate 4.5 to 1 hoax-to-fact ratio. Evaluation uses five-seed runs with McNemar and paired-bootstrap significance testing. The proposed configuration achieves an F1-macro of 0.7240, accuracy of 0.8392, ROC-AUC of 0.8083, and a Matthews correlation coefficient of 0.4510, significantly outperforming every TF-IDF baseline at the 0.05 level. Text augmentation does not provide consistent gains, which implies that imbalance handling is the more effective lever for Indonesian hoax detection.

References

1] M. F. Mridha, A. J. Keya, Md. A. Hamid, M. M. Monowar, and Md. S. Rahman, “A Comprehensive Review on Fake News Detection With Deep Learning,” IEEE Access, vol. 9, pp. 156151–156170, 2021, doi: 10.1109/ACCESS.2021.3129329.

[2] F. Koto, A. Rahimi, J. H. Lau, and T. Baldwin, “IndoLEM and IndoBERT: A Benchmark Dataset and Pre-trained Language Model for Indonesian NLP,” in Proceedings of the 28th International Conference on Computational Linguistics, Barcelona, Spain (Online): International Committee on Computational Linguistics, 2020, pp. 757–770, doi: 10.18653/v1/2020.coling-main.66.

[3] B. Wilie et al., “IndoNLU: Benchmark and Resources for Evaluating Indonesian Natural Language Understanding,” in Proceedings of the 1st Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics and the 10th International Joint Conference on Natural Language Processing, Suzhou, China: Association for Computational Linguistics, 2020, pp. 843–857, doi: 10.18653/v1/2020.aacl-main.85.

[4] M. Y. Ridho, and E. Yulianti, “From Text to Truth: Leveraging IndoBERT and Machine Learning Models for Hoax Detection in Indonesian News,” Jurnal Ilmiah Teknik Elektro Komputer dan Informatika (JITEKI), vol. 10, no. 3, pp. 544–555, 2024.

[5] D. Y. Yefferson, V. Lawijaya, and A. S. Girsang, “Hybrid model: IndoBERT and long short-term memory for detecting Indonesian hoax news,” IAES International Journal of Artificial Intelligence (IJ-AI), vol. 13, no. 2, pp. 1913–1924, 2024, doi: 10.11591/ijai.v13.i2.pp1913-1924.

[6] S. Henning, W. Beluch, A. Fraser, and A. Friedrich, “A Survey of Methods for Addressing Class Imbalance in Deep-Learning Based Natural Language Processing,” 2023, pp. 523–540, doi: 10.18653/v1/2023.eacl-main.38.

[7] M. Bayer, M.-A. Kaufhold, and C. Reuter, “A Survey on Data Augmentation for Text Classification,” ACM Computing Surveys, vol. 55, no. 7, p. 146, 2022, doi: 10.1145/3544558.

[8] D. Zha et al., “Data-centric Artificial Intelligence: A Survey,” ACM Computing Surveys, vol. 57, no. 5, 2025, doi: 10.1145/3711118.

[9] M. I. K. Sinapoy, Y. Sibaroni, and S. S. Prasetyowati, “Comparison of LSTM and IndoBERT Method in Identifying Hoax on Twitter,” Jurnal RESTI (Rekayasa Sistem dan Teknologi Informasi), vol. 7, no. 3, pp. 657–662, 2023.

[10] F. Muftie, and M. Haris, “IndoBERT Based Data Augmentation for Indonesian Text Classification,” 2023, pp. 128–132, https://ieeexplore.ieee.org/document/10250061/.

[11] J. Tian et al., “Re-embedding Difficult Samples via Mutual Information Constrained Semantically Oversampling for Imbalanced Text Classification,” 2021, doi: 10.18653/v1/2021.emnlp-main.252.

[12] K. R. M. Fernando, and C. P. Tsokos, “Dynamically Weighted Balanced Loss: Class Imbalanced Learning and Confidence Calibration of Deep Neural Networks,” IEEE Transactions on Neural Networks and Learning Systems, vol. 33, no. 7, pp. 2940–2951, 2022, doi: 10.1109/TNNLS.2020.3047335.

[13] T.-Y. Lin, P. Goyal, R. Girshick, K. He, and P. Dollár, “Focal Loss for Dense Object Detection,” in Proceedings of the IEEE International Conference on Computer Vision (ICCV 2017), 2017, pp. 2980–2988, doi: 10.1109/ICCV.2017.324.

[14] Y. Cui, M. Jia, T.-Y. Lin, Y. Song, and S. Belongie, “Class-Balanced Loss Based on Effective Number of Samples,” 2019, pp. 9268–9277, doi: 10.1109/CVPR.2019.00949.

[15] K. Cao, C. Wei, A. Gaidon, N. Arechiga, and T. Ma, “Learning Imbalanced Datasets with Label-Distribution-Aware Margin Loss,” 2019, pp. 1567–1578, https://proceedings.neurips.cc/paper/2019/hash/621461af90cadfdaf0e8d4cc25129f91-Abstract.html.

[16] J. Chen, D. Tam, C. Raffel, M. Bansal, and D. Yang, “An Empirical Survey of Data Augmentation for Limited Data Learning in NLP,” Transactions of the Association for Computational Linguistics, vol. 11, pp. 191–211, 2023, doi: 10.1162/tacl_a_00542.

[17] T. G. Dietterich, “Approximate Statistical Tests for Comparing Supervised Classification Learning Algorithms,” Neural Computation, vol. 10, no. 7, pp. 1895–1923, 1998, doi: 10.1162/089976698300017197.

[18] D. Ulmer, C. Hardmeier, and J. Frellsen, “deep-significance: Easy and Meaningful Statistical Significance Testing in the Age of Neural Networks,” 2022, doi: 10.48550/arXiv.2204.06815.

[19] J. Dodge, S. Gururangan, D. Card, R. Schwartz, and N. A. Smith, “Show Your Work: Improved Reporting of Experimental Results,” 2019, pp. 2185–2194, doi: 10.18653/v1/D19-1224.

[20] D. Chicco, and G. Jurman, “The advantages of the Matthews correlation coefficient (MCC) over F1 score and accuracy in binary classification evaluation,” BMC Genomics, vol. 21, p. 6, 2020, doi: 10.1186/s12864-019-6413-7.

Downloads

Published

21-07-2026

How to Cite

Muttaqin, R., Rosadi, M. E., Firdaus, M. I., & Agustini, D. (2026). Deteksi Hoaks Berita Berbahasa Indonesia Menggunakan IndoBERT dengan Penanganan Ketidakseimbangan Kelas Berbasis Gabungan Class Weighting dan Focal Loss. Infotek: Jurnal Informatika Dan Teknologi, 9(2), 786–797. https://doi.org/10.29408/jit.v9i2.35214

Similar Articles

<< < 1 2 3 4 5 6 7 8 9 10 > >> 

You may also start an advanced similarity search for this article.