Human-in-the-Loop LLM Assessment for Programming Education: Design and Empirical Validation
DOI:
https://doi.org/10.29408/edumatic.v10i2.35647Keywords:
automated programming assessment, empirical validation, human-ai collaboration, large language models, programming educationAbstract
Programming instructors face the challenge of providing prompt and consistent feedback, yet manual grading becomes unsustainable in large classes. While Large Language Models (LLMs) offer grading assistance, most research is based on English-language contexts and offline assessments, creating uncertainty about their dependability and the extent of human oversight needed. This study aimed to create and assess an evaluation ecosystem that integrates learning management, AI-driven task creation, LLM grading, and human review for programming courses taught in Indonesian. Employing an ADDIE-based Research and Development approach, the system was implemented for 109 students. For grading validation, instructors independently evaluated 50 assignments without access to AI predictions. The agreement was substantial, with a mean absolute error (MAE) of 4.14, a Pearson correlation of 0.986 within a 95 percent confidence interval ranging from 0.975 to 0.992, and an intraclass correlation coefficient (ICC) of 0.977. Instructor adjustments were more frequent for open-ended tasks (34.9 percent) compared to quizzes (20.2 percent). These results contributed to the development of the task-dependent human calibration (TDHC) model. The system attained a System Usability Scale (SUS) score of 88.5 and cut grading time by 87.5 percent, facilitating focused instructor review in LLM-supported programming assessments.
References
Adeoye, M. A. (2024). Revolutionizing education: Unleashing the power of the ADDIE model for effective teaching and learning. Jurnal Pendidikan Indonesia, 13(1), 202–209. https://doi.org/10.23887/jpiundiksha.v13i1.68624
Bangor, A., Kortum, P., & Miller, J. (2009). Determining what individual SUS scores mean: Adding an adjective rating scale. Journal of Usability Studies, 4(3), 114–123.
Bernik, A., Radošević, D., & Čep, A. (2025). A comparative study of large language models in programming education: Accuracy, efficiency, and feedback in student assignment grading. Applied Sciences, 15(18), 10055.
Black, P., & Wiliam, D. (1998). Assessment and classroom learning. Assessment in Education: Principles, Policy & Practice, 5(1), 7-74. https://doi.org/10.1080/0969595980050102
Branch, R. M. (2009). Instructional design: The ADDIE approach. Springer. https://doi.org/10.1007/978-0-387-09506-6
Brooke, J. (1996). SUS: A 'quick and dirty' usability scale. In P. W. Jordan, B. Thomas, B. A. Weerdmeester, & I. L. McClelland (Eds.), Usability evaluation in industry (pp. 189–194). Taylor & Francis.
Chiang, C.-H., Chen, W.-C., Kuan, C.-Y., Yang, C., & Lee, H.-Y. (2024). Large language model as an assignment evaluator: Insights, feedback, and challenges in a 1000+ student course. Computers & Education: Artificial Intelligence, 7, Article 100274. https://doi.org/10.1016/j.caeai.2024.100274
Denny, P., Leinonen, J., Prather, J., Luxton-Reilly, A., Amarouche, T., Becker, B. A., & Reeves, B. N. (2024). Prompt problems: A new programming exercise for the generative AI era. In Proceedings of the 55th ACM Technical Symposium on Computer Science Education (pp. 296–302). ACM. https://doi.org/10.1145/3626252.3630909
Messick, S. (1995). Validity of psychological assessment: Validation of inferences from persons' responses and performances as scientific inquiry into score meaning. American Psychologist, 50(9), 741-749. https://doi.org/10.1037/0003-066X.50.9.741
Mohamed, K., Yousef, M., Medhat, W., Mohamed, E. H., Khoriba, G., & Arafa, T. (2025). Hands-on analysis of using large language models for the auto evaluation of programming assignments. Information Systems, 130, Article 102523. https://doi.org/10.1016/j.is.2024.102523
Parameswara, D. A. D., Berlilana, B., & Saputro, R. E. (2026). Utilitarian vs human-centered AI acceptance: Explaining students' adoption of ChatGPT in higher education. Edumatic: Jurnal Pendidikan Informatika, 10(1), 190–199. https://doi.org/10.29408/edumatic.v10i1.34218
Poličar, P. G., Stražar, M., & Zupan, B. (2025). Automated assignment grading with large language models: Insights from a bioinformatics course. Bioinformatics, 41(Suppl. 1), i21–i29. https://doi.org/10.1093/bioinformatics/btaf216
Pressman, R. S., & Maxim, B. R. (2020). Software engineering: A practitioner's approach (9th ed.). McGraw-Hill Education.
Suria, O. (2024). A statistical analysis of System Usability Scale (SUS) evaluations in online learning platform. Journal of Information Systems and Informatics, 6(2), 992–1007. https://doi.org/10.51519/journalisi.v6i2.750
Waleska, R. F., Asnal, H., Rahmiati, R., & Gunadi, G. (2025). Aplikasi chatbot interaktif pembelajaran bahasa pemrograman PHP dengan algoritma NLP berbasis BERT. Edumatic: Jurnal Pendidikan Informatika, 9(2), 609–618. https://doi.org/10.29408/edumatic.v9i2.31427
Yan, Z., & Pastore, S. (2022). Assessing teachers’ strategies in formative assessment: The Teacher Formative Assessment Practice Scale. Journal of Psychoeducational Assessment, 40(5), 592–604. https://doi.org/10.1177/07342829221075121
Yasmine, Y. S., & Hikmawan, R. (2025). ChatGPT sebagai alat bantu dalam penulisan karya ilmiah mahasiswa: Analisis keterlibatan dan kreativitas. Edumatic: Jurnal Pendidikan Informatika, 9(1), 99–108. https://doi.org/10.29408/edumatic.v9i1.29496
Zawacki-Richter, O., Marín, V. I., Bond, M., & Gouverneur, F. (2019). Systematic review of research on artificial intelligence applications in higher education–where are the educators?. International journal of educational technology in higher education, 16(1), 39. https://doi.org/10.1186/s41239-019-0171-0
Zhai, X., Krajcik, J., & Pellegrino, J. W. (2021). On the validity of machine learning-based next generation science assessments: A validity inferential network. Journal of Science Education and Technology, 30(2), 298–312. https://doi.org/10.1007/s10956-020-09879-9
Zimmerman, B. J. (2002). Becoming a self-regulated learner: An overview. Theory Into Practice, 41(2), 64–70. https://doi.org/10.1207/s15430421tip4102_2
Downloads
Published
How to Cite
Issue
Section
License
Copyright (c) 2026 Ahmad Muyassar Ibrahim, Runal Rezkiawan

This work is licensed under a Creative Commons Attribution-ShareAlike 4.0 International License.
All articles in this journal are the sole responsibility of the authors. Edumatic: Jurnal Pendidikan Informatika can be accessed free of charge, in accordance with the Creative Commons license used.

This work is licensed under a Lisensi a Creative Commons Attribution-ShareAlike 4.0 International License.

