Closed-Source vs. Open-Weights Language Models: Accuracy, Stability, and Deployment Trade-Offs in Educational Assessment (GPT-4o vs. LLaMA 3.1)

Authors

Keywords:

Educational Technology, GPT-4o, Large Language Models (LLMs), LLaMA 3.1, Model Comparison

Abstract

Large language models (LLMs) are increasingly used to support educational assessment, yet institutional adoption requires balancing performance with governance constraints related to auditability, data control, and deployability. We compare a closed-source model (GPT-4o) and an open-weights model (LLaMA 3.1) on two expert validated assessment formats across five academic domains (Science, History, Literature, Technology, and Social Studies). Experiment 1 evaluates multiple choice question (MCQ) option selection using 148 items, five prompting strategies (Raw, Brief Instruction, Long Instruction, Chain-of-Thought, and Question-Answer Prompt Generation), and five independent repetitions per condition (7,400 total model calls) to estimate both accuracy and run-to-run stability under a strict A to E output constraint. Experiment 2 evaluates 147 one sentence short answer items under a fixed instruction prompt. In the multiple choice questions (MCQs), the accuracy and stability of the results for GPT-4o were higher compared to those of LLaMA 3.1, and the incorporation of the additional prompt structure was associated with higher stability compared to the average improvement in correctness. In the short answer questions, the performance of GPT-4o was slightly higher compared to that of LLaMA 3.1 for ROUGE-L and METEOR, and these results mostly corresponded to the results obtained for the semantic similarity questions and the evaluation rubric. The results of the study are important for understanding the accuracy-stability trade-off for the two language generators and for supporting the evaluation approach that considers the correctness of the results for MCQs, their stability for multiple runs, and the validation of the results for short answer questions using the evaluation rubric and the results for the semantic similarity questions.

Downloads

Download data is not yet available.

Author Biographies

  • Yuniar Indrihapsari, Universitas Negeri Yogyakarta

    Yuniar Indrihapsari        is an Assistant Professor at Universitas Negeri Yogyakarta (UNY), Indonesia, and a Ph.D. student at National Taiwan University of Science and Technology, Taiwan. She received her Bachelor’s degree in Electrical Engineering and Master’s degree in Information Technology from Universitas Gadjah Mada, Indonesia. Her research interests include social network analysis, e-learning, human-computer interaction, and educational technology. In this paper, she contributed to conceptualization, writing – original draft preparation, and formal analysis. She can be contacted at email: yuniar@uny.ac.id.

  • Pradana Setialana, Universitas Negeri Yogyakarta

    Pradana Setialana        is an Assistant Professor in the Department of Electronics and Informatics Engineering Education at Universitas Negeri Yogyakarta, Indonesia. He received his bachelor degree degree in Informatics Engineering Education from Universitas Negeri Yogyakarta and his master degree in Information Technology from Universitas Gadjah Mada, Indonesia. His research interests include software engineering, natural language processing, mobile and cloud computing, and database systems. In this paper, he contributed to methodology, software development, and data curation. He can be contacted at email: pradana@uny.ac.id

  • Danang Wijaya , National Central University, Taiwan

    Danang Wijaya      received his bachelor degree in Informatics Engineering Education from Universitas Negeri Yogyakarta, Indonesia, and his master degree from the International Master’s Program in Artificial Intelligence, National Central University, Taiwan. His research interests include artificial intelligence and natural language processing. In this paper, he contributed to investigation, validation, and visualization. He can be contacted at email: danangwijaya750@gmail.com

  • Satya Adhiyaksa Ardy , National Taiwan University of Science and Technology

    Satya Adhiyaksa Ardy      received his bachelor degree in Information Technology from Universitas Negeri Yogyakarta, Indonesia. He is currently a Master’s student in the Department of Electronic and Computer Engineering at National Taiwan University of Science and Technology, Taiwan. His research interests include artificial intelligence, natural language processing, and educational technology. In this paper, he contributed to writing – review & editing, resources, and project administration. He can be contacted at email: satyaadhiyaksa@gmail.com

Downloads

Published

30-06-2026

Issue

Section

Articles

How to Cite

Jati, H., Indrihapsari, Y., Romadhon, S. ., Setialana, P. ., Wijaya , D. ., & Ardy , S. A. . (2026). Closed-Source vs. Open-Weights Language Models: Accuracy, Stability, and Deployment Trade-Offs in Educational Assessment (GPT-4o vs. LLaMA 3.1). Journal of Soft Computing and Data Mining, 7(2), 58-74. https://publisher.uthm.edu.my/ojs/index.php/jscdm/article/view/24596