Research Article

Evaluating Large Language Models as Educational Tools: A Multi-Model Analysis of Performance on the Turkish Medical Specialty Examination (TUS) Ophthalmology Questions Özet:

Volume: 6 Number: 2 July 31, 2026
TR EN

Evaluating Large Language Models as Educational Tools: A Multi-Model Analysis of Performance on the Turkish Medical Specialty Examination (TUS) Ophthalmology Questions Özet:

Abstract

This study aimed to comprehensively evaluate the performance of seven leading large language models (LLMs) from the 2024-2025 period on Turkish Medical Specialty Examination (TUS) ophthalmology questions, analyzing factors such as question type, chronology, and clinical area, while critically assessing the risk of data contamination and the technology's realistic potential in medical education. A set of 210 TUS ophthalmology questions (2015-2024) was presented to seven LLMs (OpenAI o1-preview, Claude 4.5 Sonnet, GPT-4o, Gemini 2.5 Pro, DeepSeek-R1, Llama 3 70B, Command R+) using a standardized protocol. Performance was statistically compared based on overall accuracy, question type (case-based vs. knowledge-based), chronological period (old: 2015-2019 vs. new: 2020-2024), and clinical subspecialty. All responses underwent blinded, expert qualitative coding for error type and confidence. OpenAI o1-preview achieved the highest overall accuracy (95.2%). Claude 4.5 Sonnet and GPT-4o demonstrated 100% accuracy on case-based questions. A significant performance gap existed between closed-source (average 91.0%) and open-source models (average 66.1%) (p<0.001, OR=5.14). Chronological analysis revealed a significant performance drop on new questions for several models (e.g., Command R+, DeepSeek-R1), suggesting potential data contamination. Qualitative analysis showed error concentration (over 65%) in complex clinical areas like glaucoma and retina. Hallucination rates were substantially higher in open-source models (42.5% vs. 12.8%). A weak, non-significant correlation was found between model confidence and answer accuracy (rho=0.21, p=0.09). Advanced LLMs show near-expert-level textual performance on TUS questions, highlighting their potential as complementary educational tools. However, significant performance variations, vulnerabilities in nuanced clinical judgment, suspected data contamination, high hallucination rates in open-source models, and the confidence-accuracy disconnect necessitate a critical approach. This text-based performance must not be equated with real-world, visually-intensive clinical diagnostic competence. Implementation requires rigorous human oversight, AI literacy education, and ethical frameworks.

Keywords

Ethical Statement

This study involved the analysis of publicly available, de-identified examination questions with no human participants or animal subjects. Therefore, ethics committee approval and informed consent were not required. The research adhered to ethical principles for non-interventional studies.

References

  1. Antaki, F., Touma, S., Milad, D., El-Khoury, J., & Duval, R. (2023). Evaluating the Performance of ChatGPT in Ophthalmology: An Analysis of Its Successes and Shortcomings. Ophthalmology Science, *4*(4), 100424. doi.org/10.1016/j.xops.2023.100324
  2. Aygul, Y., Olucoglu, M., & Alpkocak, A. (2024). Are Large Language Models More Successful than Humans in the Turkish Medical Specialty Exam (TUS)? arXiv preprint arXiv:2408.12305. doi: 10.48550/arXiv.2408.12305
  3. Balci, A.S., Yazar, Z., Ozturk, B.T. et al. (2024). Performance of Chatgpt in ophthalmology exam; human versus AI. International Ophthalmology, *44*, 413. doi.org/10.1007/s10792-024-03353-w
  4. Bozkurt, B., Coskun, K., & Bakal, G. (2024). Building a challenging medical dataset for comparative evaluation of classifier capabilities. Computers in Biology and Medicine, *178*, 108721. doi: 10.1016/j.compbiomed.2024.108721
  5. Carbonari, V., Veltri, P., & Guzzi, P. H. (2025). Decoding Rarity: Large Language Models in the Diagnosis of Rare Diseases. arXiv preprint arXiv:2505.17065. doi: 10.48550/arXiv.2505.17065
  6. Gilson, A., Safranek, C. W., Huang, T., et al. (2023). How Does ChatGPT Perform on the United States Medical Licensing Examination? The Implications of Large Language Models for Medical Education and Knowledge Assessment. JMIR Medical Education, *9*, e45312. doi: 10.2196/45312
  7. Gokkurt Yilmaz, B. N., et al. (2025). Evaluation of the performance of different large language models on head and neck anatomy questions in the dentistry specialization exam in Turkey. Surgical and Radiologic Anatomy, *47*(1), 211. https://doi.org/10.1007/s00276-025-03723-8
  8. Greig, E. C., Duker, J. S., & Waheed, N. K. (2020). A practical guide to optical coherence tomography angiography interpretation. International Journal of Retina and Vitreous, *6*(1), 55. https://doi.org/10.1186/s40942-020-00262-9

Details

Primary Language

English

Subjects

Medical Education

Journal Section

Research Article

Publication Date

July 31, 2026

Submission Date

March 19, 2026

Acceptance Date

April 24, 2026

Published in Issue

Year 2026 Volume: 6 Number: 2

APA
Başkan, B., Evcimen, Y., Timuçin, Ö. B., Yarımağa, İ., & Sağmış, C. (2026). Evaluating Large Language Models as Educational Tools: A Multi-Model Analysis of Performance on the Turkish Medical Specialty Examination (TUS) Ophthalmology Questions Özet: Muş Alparslan Üniversitesi Sağlık Bilimleri Dergisi, 6(2), 51-65. https://doi.org/10.71255/maunsbd.1913060
AMA
1.Başkan B, Evcimen Y, Timuçin ÖB, Yarımağa İ, Sağmış C. Evaluating Large Language Models as Educational Tools: A Multi-Model Analysis of Performance on the Turkish Medical Specialty Examination (TUS) Ophthalmology Questions Özet: Muş Alparslan Üniversitesi Sağlık Bilimleri Dergisi. 2026;6(2):51-65. doi:10.71255/maunsbd.1913060
Chicago
Başkan, Burhan, Yusuf Evcimen, Özgür Bülent Timuçin, İffet Yarımağa, and Celal Sağmış. 2026. “Evaluating Large Language Models As Educational Tools: A Multi-Model Analysis of Performance on the Turkish Medical Specialty Examination (TUS) Ophthalmology Questions Özet:”. Muş Alparslan Üniversitesi Sağlık Bilimleri Dergisi 6 (2): 51-65. https://doi.org/10.71255/maunsbd.1913060.
EndNote
Başkan B, Evcimen Y, Timuçin ÖB, Yarımağa İ, Sağmış C (July 1, 2026) Evaluating Large Language Models as Educational Tools: A Multi-Model Analysis of Performance on the Turkish Medical Specialty Examination (TUS) Ophthalmology Questions Özet: Muş Alparslan Üniversitesi Sağlık Bilimleri Dergisi 6 2 51–65.
IEEE
[1]B. Başkan, Y. Evcimen, Ö. B. Timuçin, İ. Yarımağa, and C. Sağmış, “Evaluating Large Language Models as Educational Tools: A Multi-Model Analysis of Performance on the Turkish Medical Specialty Examination (TUS) Ophthalmology Questions Özet:”, Muş Alparslan Üniversitesi Sağlık Bilimleri Dergisi, vol. 6, no. 2, pp. 51–65, July 2026, doi: 10.71255/maunsbd.1913060.
ISNAD
Başkan, Burhan - Evcimen, Yusuf - Timuçin, Özgür Bülent - Yarımağa, İffet - Sağmış, Celal. “Evaluating Large Language Models As Educational Tools: A Multi-Model Analysis of Performance on the Turkish Medical Specialty Examination (TUS) Ophthalmology Questions Özet:”. Muş Alparslan Üniversitesi Sağlık Bilimleri Dergisi 6/2 (July 1, 2026): 51-65. https://doi.org/10.71255/maunsbd.1913060.
JAMA
1.Başkan B, Evcimen Y, Timuçin ÖB, Yarımağa İ, Sağmış C. Evaluating Large Language Models as Educational Tools: A Multi-Model Analysis of Performance on the Turkish Medical Specialty Examination (TUS) Ophthalmology Questions Özet: Muş Alparslan Üniversitesi Sağlık Bilimleri Dergisi. 2026;6:51–65.
MLA
Başkan, Burhan, et al. “Evaluating Large Language Models As Educational Tools: A Multi-Model Analysis of Performance on the Turkish Medical Specialty Examination (TUS) Ophthalmology Questions Özet:”. Muş Alparslan Üniversitesi Sağlık Bilimleri Dergisi, vol. 6, no. 2, July 2026, pp. 51-65, doi:10.71255/maunsbd.1913060.
Vancouver
1.Burhan Başkan, Yusuf Evcimen, Özgür Bülent Timuçin, İffet Yarımağa, Celal Sağmış. Evaluating Large Language Models as Educational Tools: A Multi-Model Analysis of Performance on the Turkish Medical Specialty Examination (TUS) Ophthalmology Questions Özet: Muş Alparslan Üniversitesi Sağlık Bilimleri Dergisi. 2026 Jul. 1;6(2):51-65. doi:10.71255/maunsbd.1913060