Research Article

Artificial Intelligence Versus Human Raters in English Writing Assessment: The Multifaceted Rasch Model

Volume: 21 Number: 51 September 26, 2026
TR EN

Artificial Intelligence Versus Human Raters in English Writing Assessment: The Multifaceted Rasch Model

Abstract

This study examines the inter-rater reliability (IRR) of three AI tools (ChatGPT 3.5, Gemini, YouChat) and four human raters in evaluating English writing skills using the multifaceted Rasch model. The evaluation focuses on the consistency and dependability of scores assigned by the AI tools compared to human judgment. The sample comprised 206 students whose responses were evaluated using a holistic scoring rubric based on the Common European Framework of Reference for Languages (CEFR) writing competencies. Findings indicate that ChatGPT 3.5 outperforms Gemini and YouChat regarding accuracy, precision, and recall, with an overall high IRR among all raters. The results suggest that AI tools can effectively supplement human raters, providing consistent and reliable assessments, which has significant implications for educational evaluations and the potential integration of AI in various scoring scenarios. The study contributes to understanding the role of AI in educational assessment and emphasizes the potential of AI tools to improve the scalability and efficiency of scoring processes. The study also discusses the importance of AI tools in providing prompt feedback, analyzing large datasets, and ensuring scoring consistency across different contexts. It also highlights the need for hybrid models that combine AI's consistency with human assessors' contextual understanding.

Keywords

References

  1. Anastasi, A. (1979). Fields of applied psychology (2nd ed.). McGraw-Hill.
  2. Brennan, R. L. (2001). Generalizability theory. Springer-Verlag
  3. Brookhart, S.M. (2010). How to assess higher-order thinking skills in your classroom. ASCD.
  4. Brookhart, S.M. (2015). Making the most of multiple choice. educational leadership, 73(1), 36 39. https://eric.ed.gov/?id=EJ1075062
  5. Brooks, V. (1980) Improving the reliability of essay marking: a survey of the literature with particular reference to the English language composition. (CSE Research Project Report 5). Leicester University.
  6. Burton, S. J., Sudweeks, R. R., Merrill, P. F., & Wood, B. (1991). How to prepare better multiple-choice test items: Guidelines for university faculty. Bringham Young University Testing Services and the Department of Instructional Science. https://testing.byu.edu/info/handbooks/betteritems.pdf.
  7. Chen, J. (2021). Refining the teacher emotion model: Evidence from a review of literature published between 1985 and 2019. Cambridge Journal of Education, 51(3), 327–357.
  8. Creswell, J. W., & Creswell, J. D. (2017). Research design: Qualitative, quantitative, and mixed methods approaches. SAGE publications.

Details

Primary Language

English

Subjects

Measurement Theories and Applications in Education and Psychology, Classroom Measurement Practices, Measurement and Evaluation in Education (Other)

Journal Section

Research Article

Publication Date

September 26, 2026

Submission Date

September 23, 2025

Acceptance Date

March 5, 2026

Published in Issue

Year 2026 Volume: 21 Number: 51

APA
Gürdil, H., Anadol, H. Ö., Salihoglu, S., Karaman, İ. H., & Çınkır, Ş. (2026). Artificial Intelligence Versus Human Raters in English Writing Assessment: The Multifaceted Rasch Model. Bayburt Eğitim Fakültesi Dergisi, 21(51), 1187-1210. https://doi.org/10.35675/befdergi.1788625
AMA
1.Gürdil H, Anadol HÖ, Salihoglu S, Karaman İH, Çınkır Ş. Artificial Intelligence Versus Human Raters in English Writing Assessment: The Multifaceted Rasch Model. Bayburt Eğitim Fakültesi Dergisi. 2026;21(51):1187-1210. doi:10.35675/befdergi.1788625
Chicago
Gürdil, Hatice, Hatice Özlem Anadol, Salih Salihoglu, İsmail Hakkı Karaman, and Şakir Çınkır. 2026. “Artificial Intelligence Versus Human Raters in English Writing Assessment: The Multifaceted Rasch Model”. Bayburt Eğitim Fakültesi Dergisi 21 (51): 1187-1210. https://doi.org/10.35675/befdergi.1788625.
EndNote
Gürdil H, Anadol HÖ, Salihoglu S, Karaman İH, Çınkır Ş (September 1, 2026) Artificial Intelligence Versus Human Raters in English Writing Assessment: The Multifaceted Rasch Model. Bayburt Eğitim Fakültesi Dergisi 21 51 1187–1210.
IEEE
[1]H. Gürdil, H. Ö. Anadol, S. Salihoglu, İ. H. Karaman, and Ş. Çınkır, “Artificial Intelligence Versus Human Raters in English Writing Assessment: The Multifaceted Rasch Model”, Bayburt Eğitim Fakültesi Dergisi, vol. 21, no. 51, pp. 1187–1210, Sept. 2026, doi: 10.35675/befdergi.1788625.
ISNAD
Gürdil, Hatice - Anadol, Hatice Özlem - Salihoglu, Salih - Karaman, İsmail Hakkı - Çınkır, Şakir. “Artificial Intelligence Versus Human Raters in English Writing Assessment: The Multifaceted Rasch Model”. Bayburt Eğitim Fakültesi Dergisi 21/51 (September 1, 2026): 1187-1210. https://doi.org/10.35675/befdergi.1788625.
JAMA
1.Gürdil H, Anadol HÖ, Salihoglu S, Karaman İH, Çınkır Ş. Artificial Intelligence Versus Human Raters in English Writing Assessment: The Multifaceted Rasch Model. Bayburt Eğitim Fakültesi Dergisi. 2026;21:1187–1210.
MLA
Gürdil, Hatice, et al. “Artificial Intelligence Versus Human Raters in English Writing Assessment: The Multifaceted Rasch Model”. Bayburt Eğitim Fakültesi Dergisi, vol. 21, no. 51, Sept. 2026, pp. 1187-10, doi:10.35675/befdergi.1788625.
Vancouver
1.Hatice Gürdil, Hatice Özlem Anadol, Salih Salihoglu, İsmail Hakkı Karaman, Şakir Çınkır. Artificial Intelligence Versus Human Raters in English Writing Assessment: The Multifaceted Rasch Model. Bayburt Eğitim Fakültesi Dergisi. 2026 Sep. 1;21(51):1187-210. doi:10.35675/befdergi.1788625