Araştırma Makalesi

The Use of Artificial Intelligence Tools in Assessing Content Validity: A Comparative Study with Human Experts

Cilt: 59 Sayı: 2 6 Ağustos 2026
PDF İndir
EN TR

The Use of Artificial Intelligence Tools in Assessing Content Validity: A Comparative Study with Human Experts

Öz

This study investigated whether artificial intelligence (AI) systems could evaluate the content validity of B1-level English reading comprehension items in a manner comparable to human experts. A 25-item multiple-choice test was developed and rated by four human experts and four AI models (GPT-4o, Gemini 2.0, Replika, and Chatfuel). To ensure methodological transparency, a standardized few-shot prompting framework was employed. Each model was accessed in a new session to avoid context drift, and identical prompts were provided to maintain consistency across evaluations. The Content Validity Ratio (CVR) and Item Content Validity Index (I-CVI) were calculated and compared using the Wilcoxon Signed-Rank Test. The results revealed no statistically significant difference between AI- and human-generated scores, suggesting that AI evaluators produced ratings similar to those of human experts. Although fine-tuning could not be performed due to the closed-source nature of the models, standardized prompting procedures enhanced reliability and replicability. These findings indicate that AI tools can complement expert judgment in content validity assessment when methodological control is ensured. The study highlights the potential of hybrid human–AI frameworks in educational measurement, while emphasizing the importance of model transparency and context management.

Anahtar Kelimeler

Kaynakça

  1. Ahmed, H. K., & Hussein, J. A. (2020). Design and implementation of a chatbot for Kurdish language speakers using Chatfuel platform. Kurdistan Journal of Applied Research, 5(2), 117–135. https://doi.org/10.24017/science.2020.2.10
  2. Almanasreh, E., Moles, R., & Chen, T. F. (2019). Evaluation of methods used for estimating content validity. Research in Social and Administrative Pharmacy, 15(2), 214–221. https://doi.org/10.1016/j.sapharm.2018.03.066
  3. Arqub, S. A., Al-Moghrabi, D., Allareddy, V., Upadhyay, M., Vaid, N., & Yadav, S. (2024). Content analysis of AI-generated (ChatGPT) responses concerning orthodontic clear aligners. The Angle Orthodontist, 94(3), 263–272. https://doi.org/10.2319/071123-484.1
  4. Ayre, C., & Scally, A. J. (2014). Critical values for Lawshe’s content validity ratio: Revisiting the original methods of calculation. Measurement and Evaluation in Counseling and Development, 47(1), 79–86. https://doi.org/10.1177/0748175613513808
  5. Cohen, R. J., Swerdlik, M. E., & Phillips, S. M. (2021). Psychological testing and assessment. McGraw Hill.
  6. Davis, L. L. (1992). Instrument review: Getting the most from a panel of experts. Applied Nursing Research, 5(4), 194–197. https://doi.org/10.1016/S0897-1897(05)80008-4
  7. Hassabis, D., & Kavukcuoglu, K. (2024, December 11). Introducing Gemini 2.0: Our new AI model for the agentic era. Google. https://blog.google/innovation-and-ai/models-and-research/google-deepmind/google-gemini-ai-update-december-2024/
  8. Fjelland, R. (2020). Why general artificial intelligence will not be realized. Humanities and Social Sciences Communications, 7(1), 1–9. https://doi.org/10.1057/s41599-020-0494-4

Ayrıntılar

Birincil Dil

İngilizce

Konular

Eğitimde Ölçme ve Değerlendirme (Diğer)

Bölüm

Araştırma Makalesi

Yayımlanma Tarihi

6 Ağustos 2026

Gönderilme Tarihi

21 Ekim 2025

Kabul Tarihi

24 Haziran 2026

Yayımlandığı Sayı

Yıl 2026 Cilt: 59 Sayı: 2

Kaynak Göster

APA
Gürdil, H., Anadol, H. Ö., & Soğuksu, Y. B. (2026). The Use of Artificial Intelligence Tools in Assessing Content Validity: A Comparative Study with Human Experts. Ankara University Journal of Faculty of Educational Sciences (JFES), 59(2), 549-590. https://doi.org/10.30964/auebfd.1807153
AMA
1.Gürdil H, Anadol HÖ, Soğuksu YB. The Use of Artificial Intelligence Tools in Assessing Content Validity: A Comparative Study with Human Experts. AÜEBFD. 2026;59(2):549-590. doi:10.30964/auebfd.1807153
Chicago
Gürdil, Hatice, Hatice Özlem Anadol, ve Yeşim Beril Soğuksu. 2026. “The Use of Artificial Intelligence Tools in Assessing Content Validity: A Comparative Study with Human Experts”. Ankara University Journal of Faculty of Educational Sciences (JFES) 59 (2): 549-90. https://doi.org/10.30964/auebfd.1807153.
EndNote
Gürdil H, Anadol HÖ, Soğuksu YB (01 Ağustos 2026) The Use of Artificial Intelligence Tools in Assessing Content Validity: A Comparative Study with Human Experts. Ankara University Journal of Faculty of Educational Sciences (JFES) 59 2 549–590.
IEEE
[1]H. Gürdil, H. Ö. Anadol, ve Y. B. Soğuksu, “The Use of Artificial Intelligence Tools in Assessing Content Validity: A Comparative Study with Human Experts”, AÜEBFD, c. 59, sy 2, ss. 549–590, Ağu. 2026, doi: 10.30964/auebfd.1807153.
ISNAD
Gürdil, Hatice - Anadol, Hatice Özlem - Soğuksu, Yeşim Beril. “The Use of Artificial Intelligence Tools in Assessing Content Validity: A Comparative Study with Human Experts”. Ankara University Journal of Faculty of Educational Sciences (JFES) 59/2 (01 Ağustos 2026): 549-590. https://doi.org/10.30964/auebfd.1807153.
JAMA
1.Gürdil H, Anadol HÖ, Soğuksu YB. The Use of Artificial Intelligence Tools in Assessing Content Validity: A Comparative Study with Human Experts. AÜEBFD. 2026;59:549–590.
MLA
Gürdil, Hatice, vd. “The Use of Artificial Intelligence Tools in Assessing Content Validity: A Comparative Study with Human Experts”. Ankara University Journal of Faculty of Educational Sciences (JFES), c. 59, sy 2, Ağustos 2026, ss. 549-90, doi:10.30964/auebfd.1807153.
Vancouver
1.Hatice Gürdil, Hatice Özlem Anadol, Yeşim Beril Soğuksu. The Use of Artificial Intelligence Tools in Assessing Content Validity: A Comparative Study with Human Experts. AÜEBFD. 01 Ağustos 2026;59(2):549-90. doi:10.30964/auebfd.1807153

Ankara Üniversitesi Eğitim Bilimleri Fakültesi Dergisi, CC BY-NC-ND 4.0 lisansını kullanmaktadır.