Analisis Perbandingan Zero-Shot dan Fine-Tuning pada LLM Untuk Automated Essay Scoring
DOI:
https://doi.org/10.30998/jrami.v7i03.1452Keywords:
Open-weight LLM, Automated Essay Scoring (AES), Large Language Model (LLM), Fine-tuning, Zero-shotAbstract
Automated Essay Scoring (AES) using Large Language Models (LLMs) is rapidly evolving, yet there remains a lack of studies comparing Zero-shot prompting and Fine-tuning approaches on open-weight LLM models. This study analyzes the comparison of these two approaches on the Qwen 2.5-1.5B-Instruct model using the ASAP-AES 2.0 dataset. The Zero-shot approach employs Role Prompting and Chain-of-Thought (CoT) techniques, while Fine-tuning uses Parameter-Efficient Fine-tuning (PEFT/LoRA). Evaluation was conducted using the Quadratic Weighted Kappa (QWK) and Root Mean Square Error (RMSE) metrics. The results show that the Zero-shot approach yields a QWK of 0.0069 with an extreme central tendency bias, while Fine-tuning yields a QWK of 0.3247 (fair agreement). These findings confirm that adapting a Fine-tuning approach is one way to produce a sufficiently accurate automatic essay grading system on small-scale open-weight models
Downloads
References
Adha, R., Syaripudin, D., Kurniawan, B. S., Mukharil Bachtiar, A., Maulana, H., Rainarli, E., & Sistem Informasi, D. (2026). Integrasi Gemini AI Berbasis Retrieval-Augmented Generation (RAG) pada Moodle untuk Penilaian Esai Otomatis dengan Pendekatan Human-in-the-Loop di Pendidikan Tinggi. In AJCSR [Academic Journal of Computer Science Research (Vol. 8, Number 1). https://doi.org/10.38101/ajcsr.v8i1.16253.
Alawadh, H. M., Meraj, T., Aldosari, L., & Tayyab Rauf, H. (2024). An Efficient Text-Mining Framework of Automatic Essay Grading Using Discourse Macrostructural and Statistical Lexical Features. SAGE Open, 14(4). https://doi.org/10.1177/21582440241300548.
AlGhamdi, E. M., Li, Y., Gašević, D., & Chen, G. (2026). Leveraging prompt-based LLMs for automated scoring and feedback generation in higher education. Computers and Education, 243. https://doi.org/10.1016/j.compedu.2025.105511.
Bouaziz, A., Tayaa, K., & Menci, A. (2025). Evaluation and Assessment of Essays in Higher Education: Practices, Challenges, and Insights. Анали Филолошког Факултета, 37(2), 195–218. https://doi.org/10.18485/analiff.2025.37.2.10.
Choi, H., Kang, M. C., Seong, J., & Huang, J. X. (2026). Exploring Zero-shot essay scoring: from feature-based to LLM-based approaches. Data Mining and Knowledge Discovery, 40(3). https://doi.org/10.1007/s10618-026-01193-z.
Diningrat, M. S. M., Mabruroh, F., Widya, H., & Nurhamidah, N. (2025). The Application of Machine Learning Algorithms in Analyzing Students’ Conceptual Error Patterns in Science Learning. Jurnal Pendidikan Dan Ilmu Fisika, 5(1), 54–65. https://doi.org/10.52434/jpif.v5i1.42586.
Faza, M. N., Purnamasari, P. D., & Ratna, A. A. P. (2025). Defying Data Scarcity: High-Performance Indonesian Short Answer Grading via Reasoning-Guided Language Model Fine-tuning. International Journal of Electrical, Computer, and Biomedical Engineering, 3(3). https://doi.org/10.62146/ijecbe.v3i3.148.
García-Varela, F., Nussbaum, M., Mendoza, M., Martínez-Troncoso, C., & Bekerman, Z. (2025). ChatGPT as a Stable and Fair Tool for Automated Essay Scoring. Education Sciences, 15(8). https://doi.org/10.3390/educsci15080946.
Huang, Y., Palermo, C., & Wilson, J. (2026). Accuracy and fairness of generative AI in automated essay scoring: Comparing GPT-4o, feature-based models, and human raters. Assessing Writing, 69. https://doi.org/10.1016/j.asw.2026.101047.
Liu, T., Ye, L., & Yan, W. (2026). A framework for evaluation of Large Language Models in essay assessment: Reliability, alignment, and causal reasoning. Computers and Education: Artificial Intelligence, 10. https://doi.org/10.1016/j.caeai.2026.100565.
Mahmud, T., Rabbi, M. F., Hossain, M. T., Talukder, A. A., & Hasan, M. K. (2025). Portfolio assessment for developing higher order thinking skills in Bangladeshi undergraduate EFL writing classes. Language Testing in Asia, 15(1). https://doi.org/10.1186/s40468-025-00404-6.
Martin, F., Kim, S., Bolliger, D. U., & DeLarm, J. (2025). Assessment Types, Strategies, and Feedback in Online Higher Education Courses in the Age of Artificial Intelligence: Perspectives of Instructional Designers. TechTrends, 69(6), 1330–1346. https://doi.org/10.1007/s11528-025-01115-8.
Mughal, N., Imran, A. S., Daudpota, S. M., Kastrati, Z., & Noor, W. (2026). Exploring potential of Large Language Models for automated essay scoring in education. Discover Artificial Intelligence, 6(1). https://doi.org/10.1007/s44163-026-01002-y.
Pulung Hendro Prastyo, Eddy Tungadi, & Shaifudin Zuhdi. (2025). Indonesian Automated Essay Scoring: A Comparative Study of Pretrained Transformer Models. Information Technology Education Journal, 120–130. https://doi.org/10.59562/intec.v4i2.8069.
Rizqi Amelia, A., & Tommy S Suyasa, P. Y. (2025). THE IMPACTS OF SELF-REGULATED LEARNING ON UNIVERSITY STUDENTS: A SCOPING REVIEW. International Journal of Application on Social Science and Humanities, 3(2), 61–72. https://doi.org/10.24912/ijassh.v3i2.35510.
Xiao, C., Ma, W., Song, Q., Xu, S. X., Zhang, K., Wang, Y., & Fu, Q. (2025). Human-AI Collaborative Essay Scoring: A Dual-Process Framework with LLMs. 15th International Conference on Learning Analytics and Knowledge, LAK 2025, 293–305. https://doi.org/10.1145/3706468.3706507.
Downloads
Published
Issue
Section
License
Copyright (c) 2026 Alexandria Felicia Seanne, Husnul Hadah, Yustika Heti Handal, Britney Levina Sukma, Budi Tjahyono (Author)

This work is licensed under a Creative Commons Attribution 4.0 International License.





