Araştırma Makalesi

Large Language Model Accuracy on Subspecialty Anaesthesiology Examination Questions

Cilt: 16 Sayı: 5 29 Eylül 2026
PDF İndir
TR EN

Large Language Model Accuracy on Subspecialty Anaesthesiology Examination Questions

Öz

Aim: In large language model (LLM) examination studies, models answer the same items repeatedly, so observations are clustered and conventional intervals overstate precision. We aimed to measure the accuracy of seven LLMs on a subspecialty-level anaesthesiology question bank and to determine which between-model and subgroup differences remain distinguishable under a cluster-respecting analysis. Materials and Methods: A 120-item, five-option multiple-choice bank in the style of the Anaesthesiology and Reanimation Subspecialty Examination, written by the authors and reviewed against current guidelines and textbooks, was administered to seven LLMs in five runs at temperature 0, with option order re-randomised per run (4,200 responses from 120 unique items). Accuracy is reported with cluster bootstrap 95% confidence intervals (CIs) resampling items (4,000 replicates) and naive Wilson intervals for comparison; subgroup analyses are exploratory. Results: Pooled accuracy was 80.1% (cluster-robust 95% CI 75.8-83.9; naive Wilson 78.8-81.3). Between-model differences were large and robust, spanning 60.8% (54.0-67.5) to 94.0% (90.3-97.0) with non-overlapping extremes. Cluster-robust intervals were wider in 25 of 26 estimates (median factor 2.1, range 0.8-4.0). Most subgroup comparisons were not distinguishable, intervals overlapping substantially: vignette (75.4%, 62.5-86.5) versus non-vignette (80.9%, 76.6-84.9) and guideline-dependent (73.0%, 64.6-80.7) versus other items (80.9%, 76.6-84.9). Only contrasts between weakest and strongest domains persisted. Between-run standard deviation was 1.26-2.80 points. Conclusion: Between-model differences and run-to-run instability are robust; commonly emphasised domain and item-characteristic differences are mostly not distinguishable once clustering is respected. Benchmarks of 120 items can separate models whose accuracies differ widely but cannot reliably localise their weaknesses.

Anahtar Kelimeler

Etik Beyan

This study did not involve human participants, animals, patient data or any identifiable personal information. It was conducted exclusively on publicly released multiple-choice examination items and answer keys published by the Measuring, Selection and Placement Centre (ÖSYM), together with the outputs generated by publicly accessible computational language models. Because no human or animal subjects and no clinical data were involved, approval by an institutional ethics committee was not required and informed consent was not applicable. The study was conducted in accordance with the principles of the Declaration of Helsinki insofar as they apply to research not involving human subjects.

Kaynakça

  1. 1. Thirunavukarasu AJ, Ting DSJ, Elangovan K, Gutierrez L, Tan TF, Ting DSW. Large language models in medicine. Nat Med 2023;29(8):1930-40.
  2. 2. Kung TH, Cheatham M, Medenilla A et al. Performance of ChatGPT on USMLE: potential for AI-assisted medical education using large language models. PLOS Digit Health 2023;2(2):e0000198.
  3. 3. Gilson A, Safranek CW, Huang T et al. How does ChatGPT perform on the United States Medical Licensing Examination (USMLE)? JMIR Med Educ 2023;9:e45312.
  4. 4. Singhal K, Azizi S, Tu T et al. Large language models encode clinical knowledge. Nature 2023;620(7972):172-80.
  5. 5. Vishwanath K, Alyakin A, Ghosh M et al. Evaluating the performance and fragility of large language models on the self-assessment for neurological surgeons. Neurosurgery 2025. doi:10.1227/neu.0000000000003878.
  6. 6. Angel MC, Rinehart JB, Cannesson MP, Baldi P. Clinical knowledge and reasoning abilities of AI large language models in anesthesiology: a comparative study on the American Board of Anesthesiology examination. Anesth Analg 2024;139(2):349-56.
  7. 7. Altermatt FR, Neyem A, Sumonte NI, Villagrán I, Mendoza M, Lacassie HJ. Evaluating the performance of large language models on the CONACEM anesthesiology certification exam: a comparison with human participants. Appl Sci 2025;15(11):6245.
  8. 8. Ronen A, Fein S, Orbach-Zinger S et al. Large language models versus human examinee performance on Israeli anesthesiology board examinations. Sci Rep 2026;16(1):14978.

Ayrıntılar

Birincil Dil

İngilizce

Konular

Ağrı, Anesteziyoloji, Yoğun Bakım

Bölüm

Araştırma Makalesi

Yayımlanma Tarihi

29 Eylül 2026

Gönderilme Tarihi

1 Ağustos 2026

Kabul Tarihi

14 Eylül 2026

Yayımlandığı Sayı

Yıl 2026 Cilt: 16 Sayı: 5

Kaynak Göster

APA
Koç, M. N., & Gülbay, S. R. (2026). Large Language Model Accuracy on Subspecialty Anaesthesiology Examination Questions. Journal of Contemporary Medicine, 16(5), 253-261. https://doi.org/10.16899/jcm.2008162
AMA
1.Koç MN, Gülbay SR. Large Language Model Accuracy on Subspecialty Anaesthesiology Examination Questions. Journal of Contemporary Medicine. 2026;16(5):253-261. doi:10.16899/jcm.2008162
Chicago
Koç, Muhammed Nezih, ve Sait Ramazan Gülbay. 2026. “Large Language Model Accuracy on Subspecialty Anaesthesiology Examination Questions”. Journal of Contemporary Medicine 16 (5): 253-61. https://doi.org/10.16899/jcm.2008162.
EndNote
Koç MN, Gülbay SR (01 Eylül 2026) Large Language Model Accuracy on Subspecialty Anaesthesiology Examination Questions. Journal of Contemporary Medicine 16 5 253–261.
IEEE
[1]M. N. Koç ve S. R. Gülbay, “Large Language Model Accuracy on Subspecialty Anaesthesiology Examination Questions”, Journal of Contemporary Medicine, c. 16, sy 5, ss. 253–261, Eyl. 2026, doi: 10.16899/jcm.2008162.
ISNAD
Koç, Muhammed Nezih - Gülbay, Sait Ramazan. “Large Language Model Accuracy on Subspecialty Anaesthesiology Examination Questions”. Journal of Contemporary Medicine 16/5 (01 Eylül 2026): 253-261. https://doi.org/10.16899/jcm.2008162.
JAMA
1.Koç MN, Gülbay SR. Large Language Model Accuracy on Subspecialty Anaesthesiology Examination Questions. Journal of Contemporary Medicine. 2026;16:253–261.
MLA
Koç, Muhammed Nezih, ve Sait Ramazan Gülbay. “Large Language Model Accuracy on Subspecialty Anaesthesiology Examination Questions”. Journal of Contemporary Medicine, c. 16, sy 5, Eylül 2026, ss. 253-61, doi:10.16899/jcm.2008162.
Vancouver
1.Muhammed Nezih Koç, Sait Ramazan Gülbay. Large Language Model Accuracy on Subspecialty Anaesthesiology Examination Questions. Journal of Contemporary Medicine. 01 Eylül 2026;16(5):253-61. doi:10.16899/jcm.2008162