Research Article

Comparative Evaluation of Four Large Language Models in Turkish Dentistry Specialization Exam

Volume: 12 Number: 4 September 19, 2025
TR EN

Comparative Evaluation of Four Large Language Models in Turkish Dentistry Specialization Exam

Abstract

Background The aim of the study is to evaluate the performance of four leading Large Language Models (LLMs) in the 2021 Dentistry Specialization Training Exam (DSE). Methods A total of 112 questions were used, including 39 questions in basic sciences and 73 questions in clinical sciences, which did not include the figures and graphs asked in the 2021 DSE. The study evaluated the performance of four LLMs: Claude-3.5 Haiku, GPT-3.5, Co-pilot, and Gemini-1.5. Results In basic sciences, Claude-3.5 Haiku and GPT-3.5 answered all questions correctly by 100%, while Gemini-1.5 answered by 94.9% and Co-pilot by 92.3%. In clinical sciences, Claude-3.5 Haiku showed an overall correct answer rate of 89%, Co-pilot 80.9%, GPT-3.5 79.7% and Gemini-1.5 65.7%. For all questions, Claude-3.5 Haiku showed a correct answer rate of 92.85%, GPT-3.5 86.6%, Co-pilot 84.8% and Gemini-1.5 75.9%. While the performance of LLMs in basic sciences was similar (p=0.134), there was a statistically significant difference between the performances of LLMs in clinical sciences and all questions (p=0.007 and p=0.005, respectively). Conclusion In all questions and clinical sciences, Claude-3.5 Haiku performed best, Gemini-1.5 performed worst, and GPT-3.5 and Co-pilot performed similarly. The 4 LLM models examined showed a higher success rate in basic sciences than in clinical sciences. The results showed that AI-based LLMs can perform well in knowledge-based questions such as basic sciences but perform poorly in questions that require knowledge as well as clinical reasoning, discussion, and interpretation, such as clinical sciences. Keywords Artificial intelligence, Dentistry, Dentistry specialization training, Large language model

Keywords

Ethical Statement

Since this study used only publicly available internet data and did not involve human participants, ethics committee approval was not required.

References

  1. 1. Dashti M, Londono J, Ghasemi S, et al. Attitudes, knowledge, and perceptions of dentists and dental students toward artificial intelligence: a systematic review. J Taibah Univ Med Sci. 2024;19(2):327-337. doi:10.1016/j.jtumed.2023.12.010
  2. 2. Chakravorty S, Aulakh BK, Shil M, Nepale M, Puthenkandathil R, Syed W. Role of Artificial Intelligence (AI) in Dentistry: A Literature Review. J Pharm Bioallied Sci. 2024;16(Suppl 1):S14-S16. doi:10.4103/JPBS. JPBS_466_23,
  3. 3. Sur J, Bose S, Khan F, Dewangan D, Sawriya E, Roul A. Knowledge, attitudes, and perceptions regarding the future of artificial intelligence in oral radiology in India: A survey. Imaging Sci Dent. 2020;50(3):193-198. doi:10.5624/ISD.2020.50.3.193
  4. 4. Eggmann F, Weiger R, Zitzmann NU, Blatz MB. Implications of large language models such as ChatGPT for dental medicine. J Esthet Restor Dent. 2023;35(7):1098-1102. doi:10.1111/JERD.13046
  5. 5. Shrivastava PK, Uppal S, Kumar G, Jha P. Role of ChatGPT in Academia: Dental Students’ Perspectives. Prim Dent J. 2024;13(1):89-90. doi:10.1177/20501684241230191,
  6. 6. Rahad K, Martin K, Amugo I, et al. ChatGPT to Enhance Learning in Dental Education at a Historically Black Medical College. Dent Res oral Heal. 2024;7(1). doi:10.26502/DROH.0069
  7. 7. Kasneci E, Sessler K, Küchemann S, et al. ChatGPT for good? On opportunities and challenges of large language models for education. Learn Individ Differ. 2023;103:102274. doi:10.1016/J.LINDIF.2023.102274
  8. 8. Thirunavukarasu AJ, Ting DSJ, Elangovan K, Gutierrez L, Tan TF, Ting DSW. Large language models in medicine. Nat Med. 2023;29(8):1930-1940. doi:10.1038/S41591-023-02448-8

Details

Primary Language

English

Subjects

Dentistry (Other)

Journal Section

Research Article

Publication Date

September 19, 2025

Submission Date

April 11, 2025

Acceptance Date

June 30, 2025

Published in Issue

Year 2025 Volume: 12 Number: 4

APA
Ekici, Ö. (2025). Comparative Evaluation of Four Large Language Models in Turkish Dentistry Specialization Exam. Selcuk Dental Journal, 12(4), 6-10. https://doi.org/10.15311/selcukdentj.1674113
AMA
1.Ekici Ö. Comparative Evaluation of Four Large Language Models in Turkish Dentistry Specialization Exam. Selcuk Dent J. 2025;12(4):6-10. doi:10.15311/selcukdentj.1674113
Chicago
Ekici, Ömer. 2025. “Comparative Evaluation of Four Large Language Models in Turkish Dentistry Specialization Exam”. Selcuk Dental Journal 12 (4): 6-10. https://doi.org/10.15311/selcukdentj.1674113.
EndNote
Ekici Ö (September 1, 2025) Comparative Evaluation of Four Large Language Models in Turkish Dentistry Specialization Exam. Selcuk Dental Journal 12 4 6–10.
IEEE
[1]Ö. Ekici, “Comparative Evaluation of Four Large Language Models in Turkish Dentistry Specialization Exam”, Selcuk Dent J, vol. 12, no. 4, pp. 6–10, Sept. 2025, doi: 10.15311/selcukdentj.1674113.
ISNAD
Ekici, Ömer. “Comparative Evaluation of Four Large Language Models in Turkish Dentistry Specialization Exam”. Selcuk Dental Journal 12/4 (September 1, 2025): 6-10. https://doi.org/10.15311/selcukdentj.1674113.
JAMA
1.Ekici Ö. Comparative Evaluation of Four Large Language Models in Turkish Dentistry Specialization Exam. Selcuk Dent J. 2025;12:6–10.
MLA
Ekici, Ömer. “Comparative Evaluation of Four Large Language Models in Turkish Dentistry Specialization Exam”. Selcuk Dental Journal, vol. 12, no. 4, Sept. 2025, pp. 6-10, doi:10.15311/selcukdentj.1674113.
Vancouver
1.Ömer Ekici. Comparative Evaluation of Four Large Language Models in Turkish Dentistry Specialization Exam. Selcuk Dent J. 2025 Sep. 1;12(4):6-10. doi:10.15311/selcukdentj.1674113

Cited By