Araştırma Makalesi

Evaluating Large Language Models as Educational Tools: A Multi-Model Analysis of Performance on the Turkish Medical Specialty Examination (TUS) Ophthalmology Questions Özet:

Cilt: 6 Sayı: 2 31 Temmuz 2026
PDF İndir
TR EN

Evaluating Large Language Models as Educational Tools: A Multi-Model Analysis of Performance on the Turkish Medical Specialty Examination (TUS) Ophthalmology Questions Özet:

Öz

This study aimed to comprehensively evaluate the performance of seven leading large language models (LLMs) from the 2024-2025 period on Turkish Medical Specialty Examination (TUS) ophthalmology questions, analyzing factors such as question type, chronology, and clinical area, while critically assessing the risk of data contamination and the technology's realistic potential in medical education. A set of 210 TUS ophthalmology questions (2015-2024) was presented to seven LLMs (OpenAI o1-preview, Claude 4.5 Sonnet, GPT-4o, Gemini 2.5 Pro, DeepSeek-R1, Llama 3 70B, Command R+) using a standardized protocol. Performance was statistically compared based on overall accuracy, question type (case-based vs. knowledge-based), chronological period (old: 2015-2019 vs. new: 2020-2024), and clinical subspecialty. All responses underwent blinded, expert qualitative coding for error type and confidence. OpenAI o1-preview achieved the highest overall accuracy (95.2%). Claude 4.5 Sonnet and GPT-4o demonstrated 100% accuracy on case-based questions. A significant performance gap existed between closed-source (average 91.0%) and open-source models (average 66.1%) (p<0.001, OR=5.14). Chronological analysis revealed a significant performance drop on new questions for several models (e.g., Command R+, DeepSeek-R1), suggesting potential data contamination. Qualitative analysis showed error concentration (over 65%) in complex clinical areas like glaucoma and retina. Hallucination rates were substantially higher in open-source models (42.5% vs. 12.8%). A weak, non-significant correlation was found between model confidence and answer accuracy (rho=0.21, p=0.09). Advanced LLMs show near-expert-level textual performance on TUS questions, highlighting their potential as complementary educational tools. However, significant performance variations, vulnerabilities in nuanced clinical judgment, suspected data contamination, high hallucination rates in open-source models, and the confidence-accuracy disconnect necessitate a critical approach. This text-based performance must not be equated with real-world, visually-intensive clinical diagnostic competence. Implementation requires rigorous human oversight, AI literacy education, and ethical frameworks.

Anahtar Kelimeler

Etik Beyan

This study involved the analysis of publicly available, de-identified examination questions with no human participants or animal subjects. Therefore, ethics committee approval and informed consent were not required. The research adhered to ethical principles for non-interventional studies.

Kaynakça

  1. Antaki, F., Touma, S., Milad, D., El-Khoury, J., & Duval, R. (2023). Evaluating the Performance of ChatGPT in Ophthalmology: An Analysis of Its Successes and Shortcomings. Ophthalmology Science, *4*(4), 100424. doi.org/10.1016/j.xops.2023.100324
  2. Aygul, Y., Olucoglu, M., & Alpkocak, A. (2024). Are Large Language Models More Successful than Humans in the Turkish Medical Specialty Exam (TUS)? arXiv preprint arXiv:2408.12305. doi: 10.48550/arXiv.2408.12305
  3. Balci, A.S., Yazar, Z., Ozturk, B.T. et al. (2024). Performance of Chatgpt in ophthalmology exam; human versus AI. International Ophthalmology, *44*, 413. doi.org/10.1007/s10792-024-03353-w
  4. Bozkurt, B., Coskun, K., & Bakal, G. (2024). Building a challenging medical dataset for comparative evaluation of classifier capabilities. Computers in Biology and Medicine, *178*, 108721. doi: 10.1016/j.compbiomed.2024.108721
  5. Carbonari, V., Veltri, P., & Guzzi, P. H. (2025). Decoding Rarity: Large Language Models in the Diagnosis of Rare Diseases. arXiv preprint arXiv:2505.17065. doi: 10.48550/arXiv.2505.17065
  6. Gilson, A., Safranek, C. W., Huang, T., et al. (2023). How Does ChatGPT Perform on the United States Medical Licensing Examination? The Implications of Large Language Models for Medical Education and Knowledge Assessment. JMIR Medical Education, *9*, e45312. doi: 10.2196/45312
  7. Gokkurt Yilmaz, B. N., et al. (2025). Evaluation of the performance of different large language models on head and neck anatomy questions in the dentistry specialization exam in Turkey. Surgical and Radiologic Anatomy, *47*(1), 211. https://doi.org/10.1007/s00276-025-03723-8
  8. Greig, E. C., Duker, J. S., & Waheed, N. K. (2020). A practical guide to optical coherence tomography angiography interpretation. International Journal of Retina and Vitreous, *6*(1), 55. https://doi.org/10.1186/s40942-020-00262-9

Ayrıntılar

Birincil Dil

İngilizce

Konular

Tıp Eğitimi

Bölüm

Araştırma Makalesi

Yayımlanma Tarihi

31 Temmuz 2026

Gönderilme Tarihi

19 Mart 2026

Kabul Tarihi

24 Nisan 2026

Yayımlandığı Sayı

Yıl 2026 Cilt: 6 Sayı: 2

Kaynak Göster

APA
Başkan, B., Evcimen, Y., Timuçin, Ö. B., Yarımağa, İ., & Sağmış, C. (2026). Evaluating Large Language Models as Educational Tools: A Multi-Model Analysis of Performance on the Turkish Medical Specialty Examination (TUS) Ophthalmology Questions Özet: Muş Alparslan Üniversitesi Sağlık Bilimleri Dergisi, 6(2), 51-65. https://doi.org/10.71255/maunsbd.1913060
AMA
1.Başkan B, Evcimen Y, Timuçin ÖB, Yarımağa İ, Sağmış C. Evaluating Large Language Models as Educational Tools: A Multi-Model Analysis of Performance on the Turkish Medical Specialty Examination (TUS) Ophthalmology Questions Özet: Muş Alparslan Üniversitesi Sağlık Bilimleri Dergisi. 2026;6(2):51-65. doi:10.71255/maunsbd.1913060
Chicago
Başkan, Burhan, Yusuf Evcimen, Özgür Bülent Timuçin, İffet Yarımağa, ve Celal Sağmış. 2026. “Evaluating Large Language Models as Educational Tools: A Multi-Model Analysis of Performance on the Turkish Medical Specialty Examination (TUS) Ophthalmology Questions Özet:”. Muş Alparslan Üniversitesi Sağlık Bilimleri Dergisi 6 (2): 51-65. https://doi.org/10.71255/maunsbd.1913060.
EndNote
Başkan B, Evcimen Y, Timuçin ÖB, Yarımağa İ, Sağmış C (01 Temmuz 2026) Evaluating Large Language Models as Educational Tools: A Multi-Model Analysis of Performance on the Turkish Medical Specialty Examination (TUS) Ophthalmology Questions Özet: Muş Alparslan Üniversitesi Sağlık Bilimleri Dergisi 6 2 51–65.
IEEE
[1]B. Başkan, Y. Evcimen, Ö. B. Timuçin, İ. Yarımağa, ve C. Sağmış, “Evaluating Large Language Models as Educational Tools: A Multi-Model Analysis of Performance on the Turkish Medical Specialty Examination (TUS) Ophthalmology Questions Özet:”, Muş Alparslan Üniversitesi Sağlık Bilimleri Dergisi, c. 6, sy 2, ss. 51–65, Tem. 2026, doi: 10.71255/maunsbd.1913060.
ISNAD
Başkan, Burhan - Evcimen, Yusuf - Timuçin, Özgür Bülent - Yarımağa, İffet - Sağmış, Celal. “Evaluating Large Language Models as Educational Tools: A Multi-Model Analysis of Performance on the Turkish Medical Specialty Examination (TUS) Ophthalmology Questions Özet:”. Muş Alparslan Üniversitesi Sağlık Bilimleri Dergisi 6/2 (01 Temmuz 2026): 51-65. https://doi.org/10.71255/maunsbd.1913060.
JAMA
1.Başkan B, Evcimen Y, Timuçin ÖB, Yarımağa İ, Sağmış C. Evaluating Large Language Models as Educational Tools: A Multi-Model Analysis of Performance on the Turkish Medical Specialty Examination (TUS) Ophthalmology Questions Özet: Muş Alparslan Üniversitesi Sağlık Bilimleri Dergisi. 2026;6:51–65.
MLA
Başkan, Burhan, vd. “Evaluating Large Language Models as Educational Tools: A Multi-Model Analysis of Performance on the Turkish Medical Specialty Examination (TUS) Ophthalmology Questions Özet:”. Muş Alparslan Üniversitesi Sağlık Bilimleri Dergisi, c. 6, sy 2, Temmuz 2026, ss. 51-65, doi:10.71255/maunsbd.1913060.
Vancouver
1.Burhan Başkan, Yusuf Evcimen, Özgür Bülent Timuçin, İffet Yarımağa, Celal Sağmış. Evaluating Large Language Models as Educational Tools: A Multi-Model Analysis of Performance on the Turkish Medical Specialty Examination (TUS) Ophthalmology Questions Özet: Muş Alparslan Üniversitesi Sağlık Bilimleri Dergisi. 01 Temmuz 2026;6(2):51-65. doi:10.71255/maunsbd.1913060