Research Article

The Use of Artificial Intelligence Tools in Assessing Content Validity: A Comparative Study with Human Experts

Volume: 59 Number: 2 August 6, 2026
EN TR

The Use of Artificial Intelligence Tools in Assessing Content Validity: A Comparative Study with Human Experts

Abstract

This study investigated whether artificial intelligence (AI) systems could evaluate the content validity of B1-level English reading comprehension items in a manner comparable to human experts. A 25-item multiple-choice test was developed and rated by four human experts and four AI models (GPT-4o, Gemini 2.0, Replika, and Chatfuel). To ensure methodological transparency, a standardized few-shot prompting framework was employed. Each model was accessed in a new session to avoid context drift, and identical prompts were provided to maintain consistency across evaluations. The Content Validity Ratio (CVR) and Item Content Validity Index (I-CVI) were calculated and compared using the Wilcoxon Signed-Rank Test. The results revealed no statistically significant difference between AI- and human-generated scores, suggesting that AI evaluators produced ratings similar to those of human experts. Although fine-tuning could not be performed due to the closed-source nature of the models, standardized prompting procedures enhanced reliability and replicability. These findings indicate that AI tools can complement expert judgment in content validity assessment when methodological control is ensured. The study highlights the potential of hybrid human–AI frameworks in educational measurement, while emphasizing the importance of model transparency and context management.

Keywords

References

  1. Ahmed, H. K., & Hussein, J. A. (2020). Design and implementation of a chatbot for Kurdish language speakers using Chatfuel platform. Kurdistan Journal of Applied Research, 5(2), 117–135. https://doi.org/10.24017/science.2020.2.10
  2. Almanasreh, E., Moles, R., & Chen, T. F. (2019). Evaluation of methods used for estimating content validity. Research in Social and Administrative Pharmacy, 15(2), 214–221. https://doi.org/10.1016/j.sapharm.2018.03.066
  3. Arqub, S. A., Al-Moghrabi, D., Allareddy, V., Upadhyay, M., Vaid, N., & Yadav, S. (2024). Content analysis of AI-generated (ChatGPT) responses concerning orthodontic clear aligners. The Angle Orthodontist, 94(3), 263–272. https://doi.org/10.2319/071123-484.1
  4. Ayre, C., & Scally, A. J. (2014). Critical values for Lawshe’s content validity ratio: Revisiting the original methods of calculation. Measurement and Evaluation in Counseling and Development, 47(1), 79–86. https://doi.org/10.1177/0748175613513808
  5. Cohen, R. J., Swerdlik, M. E., & Phillips, S. M. (2021). Psychological testing and assessment. McGraw Hill.
  6. Davis, L. L. (1992). Instrument review: Getting the most from a panel of experts. Applied Nursing Research, 5(4), 194–197. https://doi.org/10.1016/S0897-1897(05)80008-4
  7. Hassabis, D., & Kavukcuoglu, K. (2024, December 11). Introducing Gemini 2.0: Our new AI model for the agentic era. Google. https://blog.google/innovation-and-ai/models-and-research/google-deepmind/google-gemini-ai-update-december-2024/
  8. Fjelland, R. (2020). Why general artificial intelligence will not be realized. Humanities and Social Sciences Communications, 7(1), 1–9. https://doi.org/10.1057/s41599-020-0494-4

Details

Primary Language

English

Subjects

Measurement and Evaluation in Education (Other)

Journal Section

Research Article

Publication Date

August 6, 2026

Submission Date

October 21, 2025

Acceptance Date

June 24, 2026

Published in Issue

Year 2026 Volume: 59 Number: 2

APA
Gürdil, H., Anadol, H. Ö., & Soğuksu, Y. B. (2026). The Use of Artificial Intelligence Tools in Assessing Content Validity: A Comparative Study with Human Experts. Ankara University Journal of Faculty of Educational Sciences (JFES), 59(2), 549-590. https://doi.org/10.30964/auebfd.1807153
AMA
1.Gürdil H, Anadol HÖ, Soğuksu YB. The Use of Artificial Intelligence Tools in Assessing Content Validity: A Comparative Study with Human Experts. JFES. 2026;59(2):549-590. doi:10.30964/auebfd.1807153
Chicago
Gürdil, Hatice, Hatice Özlem Anadol, and Yeşim Beril Soğuksu. 2026. “The Use of Artificial Intelligence Tools in Assessing Content Validity: A Comparative Study With Human Experts”. Ankara University Journal of Faculty of Educational Sciences (JFES) 59 (2): 549-90. https://doi.org/10.30964/auebfd.1807153.
EndNote
Gürdil H, Anadol HÖ, Soğuksu YB (August 1, 2026) The Use of Artificial Intelligence Tools in Assessing Content Validity: A Comparative Study with Human Experts. Ankara University Journal of Faculty of Educational Sciences (JFES) 59 2 549–590.
IEEE
[1]H. Gürdil, H. Ö. Anadol, and Y. B. Soğuksu, “The Use of Artificial Intelligence Tools in Assessing Content Validity: A Comparative Study with Human Experts”, JFES, vol. 59, no. 2, pp. 549–590, Aug. 2026, doi: 10.30964/auebfd.1807153.
ISNAD
Gürdil, Hatice - Anadol, Hatice Özlem - Soğuksu, Yeşim Beril. “The Use of Artificial Intelligence Tools in Assessing Content Validity: A Comparative Study With Human Experts”. Ankara University Journal of Faculty of Educational Sciences (JFES) 59/2 (August 1, 2026): 549-590. https://doi.org/10.30964/auebfd.1807153.
JAMA
1.Gürdil H, Anadol HÖ, Soğuksu YB. The Use of Artificial Intelligence Tools in Assessing Content Validity: A Comparative Study with Human Experts. JFES. 2026;59:549–590.
MLA
Gürdil, Hatice, et al. “The Use of Artificial Intelligence Tools in Assessing Content Validity: A Comparative Study With Human Experts”. Ankara University Journal of Faculty of Educational Sciences (JFES), vol. 59, no. 2, Aug. 2026, pp. 549-90, doi:10.30964/auebfd.1807153.
Vancouver
1.Hatice Gürdil, Hatice Özlem Anadol, Yeşim Beril Soğuksu. The Use of Artificial Intelligence Tools in Assessing Content Validity: A Comparative Study with Human Experts. JFES. 2026 Aug. 1;59(2):549-90. doi:10.30964/auebfd.1807153

Ankara University Journal of Faculty of Educational Sciences is licensed under CC BY-NC-ND 4.0