Research Article

Comparative analysis of the performance of large language models in answering endodontics questions on the dental specialty exam

Volume: 9 Number: 3 May 19, 2026

Comparative analysis of the performance of large language models in answering endodontics questions on the dental specialty exam

Abstract

Aims: The aim of this study was to compare the accuracy of ChatGPT (GPT-5), Gemini (Gemini 2.5 Pro), and Grok (SuperGrok) on Dental Specialty Exam (DUS) endodontics questions and to examine differences by exam year and item format (text-based vs figure-based). Methods: Endodontics questions from Turkish Assessment, Selection and Placement Center (ÖSYM) DUS exams (2012-2021) were reviewed. Of 130 questions, three items officially canceled by ÖSYM were excluded, leaving 127 questions (122 text-based, 5 figure-based). Figure-based items included periapical radiographs, clinical photographs, or schematic diagrams. Questions were submitted to each model using default settings with one response per item; no tuning or repeat runs were performed. Responses were scored as correct/incorrect using the official answer key. Fisher’s exact test and, when required, Monte Carlo-corrected Fisher’s exact test were applied (p<0.05). Results: Overall accuracy was 92.1% (117/127) for Gemini, 85.8% (109/127) for ChatGPT, and 84.3% (107/127), with no significant difference among models (p=0.134). Accuracy was higher for text-based questions (ChatGPT 86.9%, Gemini 93.4%, Grok 86.1%) and lower for figure-based questions (ChatGPT 60%, Gemini 60%, Grok 40%). Within model comparisons showed a significant text-figure difference for Gemini (p=0.049) and Grok (p=0.028), but not for ChatGPT (p=0.147). The overall year-by-year analysis suggested variation for Grok, but this was not retained in pairwise comparisons. Conclusion: Large language model (LLM) based systems showed high accuracy on DUS endodontics questions; however, performance depended on item format. These systems may support exam preparation, but outputs-particularly for figure-supported questions-should be used cautiously for verification.

Keywords

Ethical Statement

This study did not require ethics committee approval because it involved neither human participants nor animals and analyzed only publicly available ÖSYM examination questions. No personal data were collected or stored.

References

  1. Turosz N, Chęcińska K, Chęciński M, Brzozowska A, Nowak Z, Sikora M. Applications of Artificial Intelligence in the analysis of dental panoramic radiographs: an overview of systematic reviews. Dentomaxillofac Radiol. 2023;52(7):20230284. doi:10.1259/dmfr.20230284
  2. Schwendicke F, Samek W, Krois J. Artificial Intelligence in dentistry: opportunities and challenges. J Dent Res. 2020;99(7):769-774. doi:10. 1177/0022034520915714
  3. Thirunavukarasu AJ, Ting DSJ, Elangovan K, Gutierrez L, Tan TF, Ting DSW. Large language models in medicine. Nat Med. 2023;29(8):1930-1940. doi:10.1038/s41591-023-02448-8
  4. Alhaidry HM, Fatani B, Alrayes JO, Almana AM, Alfhaed NK. ChatGPT in dentistry: a comprehensive review. Cureus. 2023;15(4):e38317. doi:10. 7759/cureus.38317
  5. Riedemann L, Labonne M, Gilbert S. The path forward for large language models in medicine is open. Npj Digit Med. 2024;7(1):339. doi: 10.1038/s41746-024-01344-w
  6. Dave M, Tattar R, Alafaleg R, et al. Performance of large language models (ChatGPT4-0, Grok2 and Gemini) in UK dentistry and dental hygiene and therapy assessments. Br Dent J. 2025. doi:10.1038/s41415-025-8383-2
  7. Zhou J, Li H, Chen S, et al. Large language models in biomedicine and healthcare. npj Artif Intell. 2025;1:44. doi:10.1038/s44387-025-00047-1
  8. Eggmann F, Weiger R, Zitzmann NU, Blatz MB. Implications of large language models such as ChatGPT for dental medicine. J Esthet Restor Dent. 2023;35(7):1098-1102. doi:10.1111/jerd.13046

Details

Primary Language

English

Subjects

Endodontics

Journal Section

Research Article

Publication Date

May 19, 2026

Submission Date

March 4, 2026

Acceptance Date

May 6, 2026

Published in Issue

Year 2026 Volume: 9 Number: 3

APA
Özden, G. F., & Şahin, E. N. (2026). Comparative analysis of the performance of large language models in answering endodontics questions on the dental specialty exam. Journal of Health Sciences and Medicine, 9(3), 767-773. https://doi.org/10.32322/jhsm.1902303
AMA
1.Özden GF, Şahin EN. Comparative analysis of the performance of large language models in answering endodontics questions on the dental specialty exam. J Health Sci Med / JHSM. 2026;9(3):767-773. doi:10.32322/jhsm.1902303
Chicago
Özden, Gizem Fatma, and Elyase Nur Şahin. 2026. “Comparative Analysis of the Performance of Large Language Models in Answering Endodontics Questions on the Dental Specialty Exam”. Journal of Health Sciences and Medicine 9 (3): 767-73. https://doi.org/10.32322/jhsm.1902303.
EndNote
Özden GF, Şahin EN (May 1, 2026) Comparative analysis of the performance of large language models in answering endodontics questions on the dental specialty exam. Journal of Health Sciences and Medicine 9 3 767–773.
IEEE
[1]G. F. Özden and E. N. Şahin, “Comparative analysis of the performance of large language models in answering endodontics questions on the dental specialty exam”, J Health Sci Med / JHSM, vol. 9, no. 3, pp. 767–773, May 2026, doi: 10.32322/jhsm.1902303.
ISNAD
Özden, Gizem Fatma - Şahin, Elyase Nur. “Comparative Analysis of the Performance of Large Language Models in Answering Endodontics Questions on the Dental Specialty Exam”. Journal of Health Sciences and Medicine 9/3 (May 1, 2026): 767-773. https://doi.org/10.32322/jhsm.1902303.
JAMA
1.Özden GF, Şahin EN. Comparative analysis of the performance of large language models in answering endodontics questions on the dental specialty exam. J Health Sci Med / JHSM. 2026;9:767–773.
MLA
Özden, Gizem Fatma, and Elyase Nur Şahin. “Comparative Analysis of the Performance of Large Language Models in Answering Endodontics Questions on the Dental Specialty Exam”. Journal of Health Sciences and Medicine, vol. 9, no. 3, May 2026, pp. 767-73, doi:10.32322/jhsm.1902303.
Vancouver
1.Gizem Fatma Özden, Elyase Nur Şahin. Comparative analysis of the performance of large language models in answering endodontics questions on the dental specialty exam. J Health Sci Med / JHSM. 2026 May 1;9(3):767-73. doi:10.32322/jhsm.1902303

Interuniversity Board (UAK) Equivalency: Article published in Ulakbim TR Index journal [10 POINTS], and Article published in other (excuding 1a, b, c) international indexed journal (1d) [5 POINTS].

The Directories (indexes) and Platforms we are included in are at the bottom of the page.

Note: Our journal is not WOS indexed and therefore is not classified as Q.

You can download Council of Higher Education (CoHG) [Yüksek Öğretim Kurumu (YÖK)] Criteria) decisions about predatory/questionable journals and the author's clarification text and journal charge policy from your browser. https://dergipark.org.tr/tr/journal/2316/file/4905/show







The indexes of the journal are ULAKBİM TR Dizin, ICI World of Journals, DOAJ, Directory of Research Journals Indexing (DRJI), General Impact Factor, ASOS Index, WorldCat (OCLC), MIAR, OpenAIRE, Türkiye Citation Index, Türk Medline Index, InfoBase Index, Scilit, etc.

       images?q=tbn:ANd9GcRB9r6zRLDl0Pz7om2DQkiTQXqDtuq64Eb1Qg&usqp=CAU

500px-WorldCat_logo.svg.png

atifdizini.png

logo_world_of_journals_no_margin.png

images?q=tbn%3AANd9GcTNpvUjQ4Ffc6uQBqMQrqYMR53c7bRqD9rohCINkko0Y1a_hPSn&usqp=CAU

doaj.png  

images?q=tbn:ANd9GcSpOQFsFv3RdX0lIQJC3SwkFIA-CceHin_ujli_JrqBy3A32A_Tx_oMoIZn96EcrpLwTQg&usqp=CAU

ici2.png

asos-index.png

drji.png





The platforms of the journal are Google Scholar, CrossRef (DOI), ResearchBib, Open Access, COPE, ICMJE, NCBI, ORCID, Creative Commons, etc.

COPE-logo-300x199.jpgimages?q=tbn:ANd9GcQR6_qdgvxMP9owgnYzJ1M6CS_XzR_d7orTjA&usqp=CAU

icmje_1_orig.png

cc.logo.large.png

ncbi.pngimages?q=tbn:ANd9GcRBcJw8ia8S9TI4Fun5vj3HPzEcEKIvF_jtnw&usqp=CAU

ORCID_logo.png

1*mvsP194Golg0Dmo2rjJ-oQ.jpeg


Our Journal using the DergiPark system indexed are;

Ulakbim TR Dizin,  Index Copernicus, ICI World of JournalsDirectory of Research Journals Indexing (DRJI), General Impact FactorASOS Index, OpenAIRE, MIAR,  EuroPub, WorldCat (OCLC)DOAJ,  Türkiye Citation Index, Türk Medline Index, InfoBase Index


Our Journal using the DergiPark system platforms are;

Google, Google Scholar, CrossRef (DOI), ResearchBib, ICJME, COPE, NCBI, ORCID, Creative Commons, Open Access, and etc.


Journal articles are evaluated as "Double-Blind Peer Review". 

Our journal has adopted the Open Access Policy and articles in JHSM are Open Access and fully comply with Open Access instructions. All articles in the system can be accessed and read without a journal user.  https//dergipark.org.tr/tr/pub/jhsm/page/9535

Journal charge policy   https://dergipark.org.tr/tr/pub/jhsm/page/10912

Our journal has been indexed in DOAJ as of May 18, 2020.

Our journal has been indexed in TR-Dizin as of March 12, 2021.


17873

Articles published in Journal of Health Sciences and Medicine have open access and are licensed under the Creative Commons CC BY-NC-ND 4.0 International License.