Evaluation of ChatGPT's Performance in Residency Training Progress Exams and Competency Exams in Orthopedics and Traumatology
Öz
Background: Artificial intelligence (AI) technologies have rapidly expanded into the field of medical education, offering innovative tools for training and assessment.This study aimed to evaluate the performance of the ChatGPT-3.5 language model in the “Residency Training Progress Examination” (UEGS) and the “Competency Examination” administered by the Turkish Society of Orthopedics and Traumatology (TOTBID). The objective was to determine whether ChatGPT performs comparably to orthopedic residents and whether it can achieve a passing score in the Competency Exam. Methods: A total of 2,000 UEGS and 1,000 Competency Exam questions (2012–2023, excluding 2020) were presented to ChatGPT-3.5 using standardized prompts designed within the Role–Goals–Context (RGC) framework. The model’s responses were statistically compared with those of orthopedic residents and specialists using the Mann–Whitney U and Kruskal–Wallis tests (p < 0.05). Results: ChatGPT achieved the highest accuracy in the General Orthopedics category (62%) and the lowest in Adult Reconstructive Surgery (40%). It outperformed residents only in the Spine Surgery category (p < 0.05). In the Competency Exams, ChatGPT passed four of ten exams. Conclusion: ChatGPT-3.5 demonstrated limited reliability and accuracy in orthopedic examinations and should be used cautiously as an educational support tool. Future studies involving newer multimodal versions of large language models may clarify their potential role in medical education and assessment.
Anahtar Kelimeler
Etik Beyan
Kaynakça
- Acaroğlu, E., Kahraman, S., Senköylü, A., Berk, H., Caner, H., Özkan, S., ... (2014). Core curriculum (CC) of spinal surgery: A step forward in defining our profession. Acta Orthopaedica et Traumatologica Turcica, 48(5), 475–478.
- Alessandri Bonetti, M., Giorgino, R., Gallo Afflitto, G., De Lorenzi, F., & Egro, F. M. (2024). How does ChatGPT perform on the Italian Residency Admission National Exam compared to 15,869 medical graduates? Annals of Biomedical Engineering, 52(4), 745–749.
- Aljindan, F. K., Al Qurashi, A. A., Albalawi, I. A. S., Alanazi, A. M. M., Aljuhani, H. A. M., Falah Almutairi, F., ... (2023). ChatGPT conquers the Saudi Medical Licensing Exam: Exploring the accuracy of artificial intelligence in medical knowledge assessment and implications for modern medical education. Cureus, 15(9), Article e45043.
- Atik, O. Ş. (2024). Artificial intelligence: Who must have autonomy the machine or the human? Joint Diseases and Related Surgery, 35(1), 1–2. Ayik, G., Kolac, U. C., Aksoy, T., Yilmaz, A., Sili, M. V., Tokgozoglu, M., ... (2025). Exploring the role of artificial intelligence in Turkish orthopedic progression exams. Acta Orthopaedica et Traumatologica Turcica, 59(1), 18–26.
- Benli, İ., & Acaroğlu, E. (2011). Türk Ortopedi ve Travmatoloji Birliği Derneği (TOTBİD) Türk Ortopedi ve Travmatoloji Eğitim Konseyi Yeterlik Sınavları. Acta Orthopaedica et Traumatologica Turcica, 45(2). https://dergipark.org.tr/en/download/article-file/169969
- Gönen, D. E. (2013). 2012-2013 TOTBİD-TOTEK Uzmanlık Eğitimi Gelişim Sınavı Raporu (UEGS). Türk Ortopedi ve Travmatoloji Birliği Derneği. https://totbid.org.tr/uploads/files/uegs_2013_rapor.pdf
- Haenlein, M., & Kaplan, A. (2019). A brief history of artificial intelligence: On the past, present, and future of artificial intelligence. California Management Review, 61(4), 5–14.
- Huang, Y., Gomaa, A., Semrau, S., Haderlein, M., Lettmaier, S., Weissmann, T., ... (2023). Benchmarking ChatGPT-4 on a radiation oncology in-training exam and Red Journal Gray Zone cases: Potentials and challenges for AI-assisted medical education and decision making in radiation oncology. Frontiers in Oncology, 13, Article 1265024.
Ayrıntılar
Birincil Dil
İngilizce
Konular
Ortopedi
Bölüm
Araştırma Makalesi
Yayımlanma Tarihi
30 Mart 2026
Gönderilme Tarihi
6 Şubat 2026
Kabul Tarihi
4 Mart 2026
Yayımlandığı Sayı
Yıl 2026 Cilt: 2 Sayı: 1