Rethinking Automated Scoring in the Era of Artificial Intelligence: Emerging Directions for AI-Based Assessment
Abstract
The rapid development of generative artificial intelligence (AI), particularly large language models and multimodal models, is reshaping research on automated scoring in educational assessment. While previous research has often focused on the extent to which AI-generated scores agree with those assigned by human raters, agreement alone may not provide sufficient evidence for the quality or defensibility of an automated scoring system. This editorial argues that next-generation AI-based scoring research should move beyond model-level performance comparisons and examine the broader conditions under which AI-generated scores are produced, evaluated and used. It first discusses key design dimensions of AI-based scoring systems, including prompt design, rubric design, example-based calibration, model selection, scoring strategies and continuous performance monitoring. It then identifies emerging research directions involving evidence-grounded scoring, uncertainty and confidence calibration, process- and reasoning-aware evaluation, multimodal evidence, retrieval-augmented scoring, fairness and human–AI orchestration. Particular attention is given to iterative multi-agent and deliberative scoring, in which AI systems can review, critique and reconsider previous scoring decisions rather than independently producing a final score. The editorial concludes by proposing a broader research agenda in which automated scoring is evaluated not only in terms of agreement or accuracy, but also validity, reliability, fairness, uncertainty, transparency and the conditions under which human oversight should be incorporated.
Keywords
References
- Boduroglu, E., & Yigiter, M. S. (2026). Artificial intelligence scoring attitudes: scale development and validation. Education and Information Technologies, 31(3), 701–726. https://doi.org/10.1007/s10639-025-13836-7
- Boduroğlu, E., Koç, O., & Yiğiter, M. S. (2023). Madde Güçlüklerinin Tahmin Edilmesinde Uzman Görüşleri ve ChatGPT Performansının Karşılaştırılması / Comparison of Expert Opinions and ChatGPT Performance in Predicting Item Difficulties. Disiplinlerarası Eğitim Araştırmaları Dergisi, 7(15), 202–210. https://doi.org/10.57135/jier.1296255
- Ercikan, K., & McCaffrey, D. F. (2022). Optimizing implementation of Artificial‐intelligence‐based automated scoring: An evidence centered design approach for designing assessments for AI‐based scoring. Journal of Educational Measurement, 59(3), 272–287. https://doi.org/10.1111/jedm.12332
- Fianu, E., Amankwah-Sarfo, F., Ofori, P., Amoako, J. K., & Sumani, H. (2026). From traditional machine learning models to large language models: A systematic literature review of automated essay scoring. SN Computer Science, 7(5). https://doi.org/10.1007/s42979-026-05028-y
- Haudek, K. C., & Zhai, X. (2024). Examining the effect of assessment construct characteristics on machine learning scoring of scientific argumentation. International Journal of Artificial Intelligence in Education, 34(4), 1482–1509. https://doi.org/10.1007/s40593-023-00385-8
- Humphry, S. M., & Heldsinger, S. A. (2014). Common structural design features of rubrics may represent a threat to validity. Educational Researcher (Washington, D.C.: 1972), 43(5), 253–263. https://doi.org/10.3102/0013189x14542154
- Jang, J., Moon, A., Jung, M., Kim, Y., & Lee, S. J. (2025). LLM agents at the roundtable: A multi-perspective and dialectical reasoning framework for essay scoring. Findings of the Association for Computational Linguistics: EMNLP 2025, 19674–19687.
- Jung, J. Y., Tyack, L., & von Davier, M. (2024). Combining machine translation and automated scoring in international large-scale assessments. Large-Scale Assessments in Education, 12(1). https://doi.org/10.1186/s40536-024-00199-7
Details
Primary Language
English
Subjects
Testing, Assessment and Psychometrics (Other)
Journal Section
Editorial
Authors
Publication Date
October 1, 2026
Submission Date
September 1, 2026
Acceptance Date
September 30, 2026
Published in Issue
Year 2026 Volume: 17 Number: 3