Human-in-the-Loop LLM Assessment for Programming Education: Design and Empirical Validation

Authors

DOI:

https://doi.org/10.29408/edumatic.v10i2.35647

Keywords:

automated programming assessment, empirical validation, human-ai collaboration, large language models, programming education

Abstract

Programming instructors face the challenge of providing prompt and consistent feedback, yet manual grading becomes unsustainable in large classes. While Large Language Models (LLMs) offer grading assistance, most research is based on English-language contexts and offline assessments, creating uncertainty about their dependability and the extent of human oversight needed. This study aimed to create and assess an evaluation ecosystem that integrates learning management, AI-driven task creation, LLM grading, and human review for programming courses taught in Indonesian. Employing an ADDIE-based Research and Development approach, the system was implemented for 109 students. For grading validation, instructors independently evaluated 50 assignments without access to AI predictions. The agreement was substantial, with a mean absolute error (MAE) of 4.14, a Pearson correlation of 0.986 within a 95 percent confidence interval ranging from 0.975 to 0.992, and an intraclass correlation coefficient (ICC) of 0.977. Instructor adjustments were more frequent for open-ended tasks (34.9 percent) compared to quizzes (20.2 percent). These results contributed to the development of the task-dependent human calibration (TDHC) model. The system attained a System Usability Scale (SUS) score of 88.5 and cut grading time by 87.5 percent, facilitating focused instructor review in LLM-supported programming assessments.

References

Adeoye, M. A. (2024). Revolutionizing education: Unleashing the power of the ADDIE model for effective teaching and learning. Jurnal Pendidikan Indonesia, 13(1), 202–209. https://doi.org/10.23887/jpiundiksha.v13i1.68624

Bangor, A., Kortum, P., & Miller, J. (2009). Determining what individual SUS scores mean: Adding an adjective rating scale. Journal of Usability Studies, 4(3), 114–123.

Bernik, A., Radošević, D., & Čep, A. (2025). A comparative study of large language models in programming education: Accuracy, efficiency, and feedback in student assignment grading. Applied Sciences, 15(18), 10055.

Black, P., & Wiliam, D. (1998). Assessment and classroom learning. Assessment in Education: Principles, Policy & Practice, 5(1), 7-74. https://doi.org/10.1080/0969595980050102

Branch, R. M. (2009). Instructional design: The ADDIE approach. Springer. https://doi.org/10.1007/978-0-387-09506-6

Brooke, J. (1996). SUS: A 'quick and dirty' usability scale. In P. W. Jordan, B. Thomas, B. A. Weerdmeester, & I. L. McClelland (Eds.), Usability evaluation in industry (pp. 189–194). Taylor & Francis.

Chiang, C.-H., Chen, W.-C., Kuan, C.-Y., Yang, C., & Lee, H.-Y. (2024). Large language model as an assignment evaluator: Insights, feedback, and challenges in a 1000+ student course. Computers & Education: Artificial Intelligence, 7, Article 100274. https://doi.org/10.1016/j.caeai.2024.100274

Denny, P., Leinonen, J., Prather, J., Luxton-Reilly, A., Amarouche, T., Becker, B. A., & Reeves, B. N. (2024). Prompt problems: A new programming exercise for the generative AI era. In Proceedings of the 55th ACM Technical Symposium on Computer Science Education (pp. 296–302). ACM. https://doi.org/10.1145/3626252.3630909

Messick, S. (1995). Validity of psychological assessment: Validation of inferences from persons' responses and performances as scientific inquiry into score meaning. American Psychologist, 50(9), 741-749. https://doi.org/10.1037/0003-066X.50.9.741

Mohamed, K., Yousef, M., Medhat, W., Mohamed, E. H., Khoriba, G., & Arafa, T. (2025). Hands-on analysis of using large language models for the auto evaluation of programming assignments. Information Systems, 130, Article 102523. https://doi.org/10.1016/j.is.2024.102523

Parameswara, D. A. D., Berlilana, B., & Saputro, R. E. (2026). Utilitarian vs human-centered AI acceptance: Explaining students' adoption of ChatGPT in higher education. Edumatic: Jurnal Pendidikan Informatika, 10(1), 190–199. https://doi.org/10.29408/edumatic.v10i1.34218

Poličar, P. G., Stražar, M., & Zupan, B. (2025). Automated assignment grading with large language models: Insights from a bioinformatics course. Bioinformatics, 41(Suppl. 1), i21–i29. https://doi.org/10.1093/bioinformatics/btaf216

Pressman, R. S., & Maxim, B. R. (2020). Software engineering: A practitioner's approach (9th ed.). McGraw-Hill Education.

Suria, O. (2024). A statistical analysis of System Usability Scale (SUS) evaluations in online learning platform. Journal of Information Systems and Informatics, 6(2), 992–1007. https://doi.org/10.51519/journalisi.v6i2.750

Waleska, R. F., Asnal, H., Rahmiati, R., & Gunadi, G. (2025). Aplikasi chatbot interaktif pembelajaran bahasa pemrograman PHP dengan algoritma NLP berbasis BERT. Edumatic: Jurnal Pendidikan Informatika, 9(2), 609–618. https://doi.org/10.29408/edumatic.v9i2.31427

Yan, Z., & Pastore, S. (2022). Assessing teachers’ strategies in formative assessment: The Teacher Formative Assessment Practice Scale. Journal of Psychoeducational Assessment, 40(5), 592–604. https://doi.org/10.1177/07342829221075121

Yasmine, Y. S., & Hikmawan, R. (2025). ChatGPT sebagai alat bantu dalam penulisan karya ilmiah mahasiswa: Analisis keterlibatan dan kreativitas. Edumatic: Jurnal Pendidikan Informatika, 9(1), 99–108. https://doi.org/10.29408/edumatic.v9i1.29496

Zawacki-Richter, O., Marín, V. I., Bond, M., & Gouverneur, F. (2019). Systematic review of research on artificial intelligence applications in higher education–where are the educators?. International journal of educational technology in higher education, 16(1), 39. https://doi.org/10.1186/s41239-019-0171-0

Zhai, X., Krajcik, J., & Pellegrino, J. W. (2021). On the validity of machine learning-based next generation science assessments: A validity inferential network. Journal of Science Education and Technology, 30(2), 298–312. https://doi.org/10.1007/s10956-020-09879-9

Zimmerman, B. J. (2002). Becoming a self-regulated learner: An overview. Theory Into Practice, 41(2), 64–70. https://doi.org/10.1207/s15430421tip4102_2

Downloads

Published

2026-08-25

How to Cite

Ibrahim, A. M., & Rezkiawan, R. (2026). Human-in-the-Loop LLM Assessment for Programming Education: Design and Empirical Validation. Edumatic: Jurnal Pendidikan Informatika, 10(2), 420–429. https://doi.org/10.29408/edumatic.v10i2.35647