EN
Optimizing AI-based assessment in history education: The impact of prompt engineering on scoring performance in a morphologically rich language
Abstract
This study investigates the effectiveness and prompt sensitivity of leading large language model (LLM)-based AI tools (ChatGPT, Gemini, Claude, Deepseek) in grading open-ended exam questions in an undergraduate history course (Atatürk’s Principles and History of Revolution), compared to human evaluators. A comprehensive dataset comprising 72 distinct open-ended responses (collected from 24 students) was scored by both the course instructor and seven AI models using five different prompts with varying levels of detail. These prompts ranged from a basic zero-shot instruction to progressively more structured designs: expected topic headings, a weighted criterion-based rubric, partial coverage with language proficiency, and relative (norm-referenced) scoring. Correlation and statistical analyses (paired samples t-test, Wilcoxon signed-rank test) revealed that while models showed low agreement with the human evaluator when using unstructured, basic prompts (zero-shot), they achieved high agreement (r > .80) when provided with structured prompts and explicit rubrics, particularly in the cases of Gemini and Claude. The findings highlight that for effective AI-based assessment, prompt design and the definition of criteria are more critical than model selection in mitigating "generosity bias." Generosity bias here denotes the models’ systematic tendency to assign higher scores than the human rater; the qualitative evidence indicates that it stems primarily from the models rewarding fluent, lengthy, and well-structured answers even when their factual content is incomplete, a tendency that explicit rubrics substantially reduced. Designed as an exploratory case study, these results provide significant empirical evidence regarding the optimization of AI as an assistive assessment tool, specifically within the context of Turkish, a morphologically rich language.
Keywords
Ethical Statement
Declaration of Conflicting Interests and Ethics
The authors declare no conflict of interest. This research study complies with research publishing ethics. The scientific and legal responsibility for manuscripts published in IJATE belongs to the author(s).
Ethics Committee Number: Bolu Abant Izzet Baysal University, Human Research Ethics Committee, 2025/247.
References
- American Educational Research Association, American Psychological Association, & National Council on Measurement in Education. (2014). Standards for educational and psychological testing. Washington, DC: American Educational Research Association.
- Attali, Y., & Burstein, J. (2006). Automated essay scoring with e-rater V.2. The Journal of Technology, Learning and Assessment, 4(3).
- Badger, E., & Thomas, B. (1992). Open-ended questions in reading. Practical Assessment, Research & Evaluation, 3(4), 1991-1993.
- Bektaş, M., & Kudubeş, A. A. (2014). Bir ölçme ve değerlendirme aracı olarak yazılı sınavlar. Dokuz Eylül Üniversitesi Hemşirelik Fakültesi Elektronik Dergisi, 7(4), 330-336.
- Brucks, M., & Toubia, O. (2025). Prompt architecture induces methodological artifacts in large language models. PloS one, 20(4), e0319159. https://doi.org/10.1371/journal.pone.0319159
- Brown, H. D. (2004). Language assessment, principles and classroom practices. USA: Longman.
- Burrows, S., Gurevych, I., & Stein, B. (2015). The eras and trends of automatic short answer grading. International Journal of Artificial Intelligence in Education, 25(1), 60-117. https://doi.org/10.1007/s40593-014-0026-8
- Capdehourat, G., Amigo, I., Lorenzo, B., & Trigo, J. (2025). On the effectiveness of LLMs for automatic grading of open-ended questions in Spanish. arXiv preprint arXiv:2503.18072. https://doi.org/10.48550/arXiv.2503.18072
Details
Primary Language
English
Subjects
Educational Technology and Computing
Journal Section
Research Article
Publication Date
September 30, 2026
Submission Date
February 5, 2026
Acceptance Date
August 19, 2026
Published in Issue
Year 2026 Volume: 9 Number: 3
APA
Özdemir, Y., & Aşçi, E. (2026). Optimizing AI-based assessment in history education: The impact of prompt engineering on scoring performance in a morphologically rich language. Journal of Educational Technology and Online Learning, 9(3), 374-396. https://doi.org/10.31681/jetol.1882892
AMA
1.Özdemir Y, Aşçi E. Optimizing AI-based assessment in history education: The impact of prompt engineering on scoring performance in a morphologically rich language. JETOL. 2026;9(3):374-396. doi:10.31681/jetol.1882892
Chicago
Özdemir, Yunus, and Emine Aşçi. 2026. “Optimizing AI-Based Assessment in History Education: The Impact of Prompt Engineering on Scoring Performance in a Morphologically Rich Language”. Journal of Educational Technology and Online Learning 9 (3): 374-96. https://doi.org/10.31681/jetol.1882892.
EndNote
Özdemir Y, Aşçi E (September 1, 2026) Optimizing AI-based assessment in history education: The impact of prompt engineering on scoring performance in a morphologically rich language. Journal of Educational Technology and Online Learning 9 3 374–396.
IEEE
[1]Y. Özdemir and E. Aşçi, “Optimizing AI-based assessment in history education: The impact of prompt engineering on scoring performance in a morphologically rich language”, JETOL, vol. 9, no. 3, pp. 374–396, Sept. 2026, doi: 10.31681/jetol.1882892.
ISNAD
Özdemir, Yunus - Aşçi, Emine. “Optimizing AI-Based Assessment in History Education: The Impact of Prompt Engineering on Scoring Performance in a Morphologically Rich Language”. Journal of Educational Technology and Online Learning 9/3 (September 1, 2026): 374-396. https://doi.org/10.31681/jetol.1882892.
JAMA
1.Özdemir Y, Aşçi E. Optimizing AI-based assessment in history education: The impact of prompt engineering on scoring performance in a morphologically rich language. JETOL. 2026;9:374–396.
MLA
Özdemir, Yunus, and Emine Aşçi. “Optimizing AI-Based Assessment in History Education: The Impact of Prompt Engineering on Scoring Performance in a Morphologically Rich Language”. Journal of Educational Technology and Online Learning, vol. 9, no. 3, Sept. 2026, pp. 374-96, doi:10.31681/jetol.1882892.
Vancouver
1.Yunus Özdemir, Emine Aşçi. Optimizing AI-based assessment in history education: The impact of prompt engineering on scoring performance in a morphologically rich language. JETOL. 2026 Sep. 1;9(3):374-96. doi:10.31681/jetol.1882892