Research Article

Diagnostic Accuracy of ChatGPT in Shoulder Fractures: A Comparative Study with an Emergency Medicine Specialist and a Radiologist

Volume: 79 Number: 2 June 30, 2026
TR EN

Diagnostic Accuracy of ChatGPT in Shoulder Fractures: A Comparative Study with an Emergency Medicine Specialist and a Radiologist

Abstract

Background Chat Generative Pre-Trained Transformer (ChatGPT) is a new large language model capable of simulating real-life conversations and providing diagnostic support. However, its performance in detecting fractures in shoulder traumas has not been investigated. Objective To evaluate the diagnostic accuracy of ChatGPT in detecting shoulder fractures and to compare its performance with an emergency medicine specialist and a radiologist. Methods This retrospective study included 197 patients who underwent both shoulder radiography and computed tomography (CT) between September 2023 and July 2025. One emergency medicine specialist, one radiologist, and ChatGPT models (4o and 4.5) independently and blindly evaluated anonymized radiographs to determine the presence and location of fractures. CT was accepted as the gold standard. Diagnostic performance was assessed using sensitivity, specificity, accuracy, and area under the curve (AUC) values. Results Among 197 patients, 74 (37.56%) had fractures, most commonly involving the humerus. The radiologist demonstrated the highest diagnostic performance (AUC: 0.903; Accuracy: 90.86%), followed by the emergency physician (AUC: 0.824; Accuracy: 82.74%). ChatGPT-4o (AUC: 0.641; Accuracy: 61.93%) and ChatGPT-4.5 (AUC: 0.626; Accuracy: 57.36%) showed statistically significant but weak-to-moderate diagnostic contribution, with high sensitivity but poor specificity. Both models demonstrated stable intra-rater agreement. Conclusion ChatGPT models showed limited performance compared with clinicians in diagnosing shoulder fractures. While they cannot replace clinical expertise, they may serve as supportive tools in decision-making.

Keywords

Shoulder trauma, Fracture, X-Ray, ChatGPT, Radiology

Ethical Statement

Approval from the Gaziantep City Hospital Coordinatorship of Local Ethics Comittee for Non-Interventional Clinical Researches was obtained during the meeting held on 18/06/2025, under decision number 212/2025.

References

  1. Baker HP, Dwyer E, Kalidoss S, et al. ChatGPT's Ability to Assist with Clinical Documentation: A Randomized Controlled Trial. J Am Acad Orthop Surg. 2024;32(3):123-129. https://doi.org/10.5435/jaaos-d-23-00474
  2. Chalhoub R, Mouawad A, Aoun M, et al. Will ChatGPT be Able to Replace a Spine Surgeon in the Clinical Setting? World Neurosurg. 2024;185:e648-e652. https://doi.org/10.1016/j.wneu.2024.02.101
  3. Liu S, Wright AP, Patterson BL, et al. Using AI-generated suggestions from ChatGPT to optimize clinical decision support. J Am Med Inform Assoc. 2023;30(7):1237-1245. https://doi.org/10.1093/jamia/ocad072
  4. Casey JC, Dworkin M, Winschel J, et al. ChatGPT: A concise Google alternative for people seeking accurate and comprehensive carpal tunnel syndrome information. Hand Surg Rehabil. 2024;43(5):101757. https://doi.org/10.1016/j.hansur.2024.101757
  5. Horiuchi D, Tatekawa H, Oura T, et al. ChatGPT's diagnostic performance based on textual vs. visual information compared to radiologists' diagnostic performance in musculoskeletal radiology. Eur Radiol. 2025;35(1):506-516. https://doi.org/10.1007/s00330-024-10902-5
  6. Horiuchi D, Tatekawa H, Shimono T, et al. Accuracy of ChatGPT generated diagnosis from patient's medical history and imaging findings in neuroradiology cases. Neuroradiology. 2024;66(1):73-79. https://doi.org/10.1007/s00234-023-03252-4
  7. De Angelis L, Baglivo F, Arzilli G, et al. ChatGPT and the rise of large language models: the new AI-driven infodemic threat in public health. Front Public Health. 2023;11:1166120. https://doi.org/10.3389/fpubh.2023.1166120
  8. Kachman MM, Brennan I, Oskvarek JJ, et al. How artificial intelligence could transform emergency care. Am J Emerg Med. 2024;81:40-46. https://doi.org/10.1016/j.ajem.2024.04.024
  9. Meral G, Ates S, Gunay S, et al. Comparative analysis of ChatGPT, Gemini and emergency medicine specialist in ESI triage assessment. Am J Emerg Med. 2024;81:146-150. https://doi.org/10.1016/j.ajem.2024.05.001
  10. Pasli S, Sahin AS, Beser MF, et al. Assessing the precision of artificial intelligence in ED triage decisions: Insights from a study with ChatGPT. Am J Emerg Med. 2024;78:170-175. https://doi.org/10.1016/j.ajem.2024.01.037
APA
Konukoğlu, O., Kaya, M., Eseoğlu Pekcan, Ş., & Günaydin, İ. (2026). Diagnostic Accuracy of ChatGPT in Shoulder Fractures: A Comparative Study with an Emergency Medicine Specialist and a Radiologist. Ankara Üniversitesi Tıp Fakültesi Mecmuası, 79(2), 216-225. https://doi.org/10.65092/autfm.1878887
AMA
1.Konukoğlu O, Kaya M, Eseoğlu Pekcan Ş, Günaydin İ. Diagnostic Accuracy of ChatGPT in Shoulder Fractures: A Comparative Study with an Emergency Medicine Specialist and a Radiologist. Ankara Üniversitesi Tıp Fakültesi Mecmuası. 2026;79(2):216-225. doi:10.65092/autfm.1878887
Chicago
Konukoğlu, Osman, Murat Kaya, Şeyma Eseoğlu Pekcan, and İsa Günaydin. 2026. “Diagnostic Accuracy of ChatGPT in Shoulder Fractures: A Comparative Study With an Emergency Medicine Specialist and a Radiologist”. Ankara Üniversitesi Tıp Fakültesi Mecmuası 79 (2): 216-25. https://doi.org/10.65092/autfm.1878887.
EndNote
Konukoğlu O, Kaya M, Eseoğlu Pekcan Ş, Günaydin İ (June 1, 2026) Diagnostic Accuracy of ChatGPT in Shoulder Fractures: A Comparative Study with an Emergency Medicine Specialist and a Radiologist. Ankara Üniversitesi Tıp Fakültesi Mecmuası 79 2 216–225.
IEEE
[1]O. Konukoğlu, M. Kaya, Ş. Eseoğlu Pekcan, and İ. Günaydin, “Diagnostic Accuracy of ChatGPT in Shoulder Fractures: A Comparative Study with an Emergency Medicine Specialist and a Radiologist”, Ankara Üniversitesi Tıp Fakültesi Mecmuası, vol. 79, no. 2, pp. 216–225, June 2026, doi: 10.65092/autfm.1878887.
ISNAD
Konukoğlu, Osman - Kaya, Murat - Eseoğlu Pekcan, Şeyma - Günaydin, İsa. “Diagnostic Accuracy of ChatGPT in Shoulder Fractures: A Comparative Study With an Emergency Medicine Specialist and a Radiologist”. Ankara Üniversitesi Tıp Fakültesi Mecmuası 79/2 (June 1, 2026): 216-225. https://doi.org/10.65092/autfm.1878887.
JAMA
1.Konukoğlu O, Kaya M, Eseoğlu Pekcan Ş, Günaydin İ. Diagnostic Accuracy of ChatGPT in Shoulder Fractures: A Comparative Study with an Emergency Medicine Specialist and a Radiologist. Ankara Üniversitesi Tıp Fakültesi Mecmuası. 2026;79:216–225.
MLA
Konukoğlu, Osman, et al. “Diagnostic Accuracy of ChatGPT in Shoulder Fractures: A Comparative Study With an Emergency Medicine Specialist and a Radiologist”. Ankara Üniversitesi Tıp Fakültesi Mecmuası, vol. 79, no. 2, June 2026, pp. 216-25, doi:10.65092/autfm.1878887.
Vancouver
1.Osman Konukoğlu, Murat Kaya, Şeyma Eseoğlu Pekcan, İsa Günaydin. Diagnostic Accuracy of ChatGPT in Shoulder Fractures: A Comparative Study with an Emergency Medicine Specialist and a Radiologist. Ankara Üniversitesi Tıp Fakültesi Mecmuası. 2026 Jun. 1;79(2):216-25. doi:10.65092/autfm.1878887