Patient Education in Hand Surgery: A Blinded Comparison of GPT-5.5 and Claude Opus 4.8
Öz
Patients increasingly consult conversational artificial intelligence about hand and wrist conditions, but the performance of proprietary large language models remains unclear. We compared the reliability, overall quality, and readability of patient-facing texts generated by GPT-5.5 and Claude Opus 4.8. Thirty hand surgery topics were selected from orthopedic education resources and the authors' outpatient practice. Each model received the same prompt and generated one response per topic (60 texts). Two fellowship-trained, certified hand surgeons independently evaluated source-masked outputs. DISCERN assessed reliability and treatment information, the Global Quality Score (GQS) overall quality, and the Flesch-Kincaid Grade Level (FKGL) and Flesch Reading Ease (FRE) readability. Platforms were compared with the Mann-Whitney U test and interrater agreement with the intraclass correlation coefficient (ICC). Median DISCERN scores were 66.5 for GPT-5.5 and 63.5 for Claude Opus 4.8 (p = 0.143); median GQS was 4.0 for both (p = 0.389). Claude Opus 4.8 had a lower median FKGL (8.75 vs 10.4; p < 0.001) and higher median FRE (50.35 vs 43.6; p = 0.009). ICCs (95% CIs) were 0.755 (0.589-0.853) for DISCERN and 0.676 (0.458-0.807) for GQS. Expert-rated reliability and overall quality did not differ significantly, whereas Claude Opus 4.8 produced more favorable formula-based readability scores; median FKGL for both models nevertheless remained above the recommended sixth-grade level. Because FKGL and FRE quantify surface linguistic complexity rather than understanding, these findings describe written outputs only; they do not establish better patient comprehension, clinical effectiveness, factual equivalence, or safety.
Anahtar Kelimeler
- Artificial Intelligence
- Large Language Models
- Patient Education as Topic
- Health Literacy
- Hand Surgery
Destekleyen Kurum
Proje Numarası
Etik Beyan
Teşekkür
Kaynakça
- 1. Ozkan S, Mellema JJ, Nazzal A, Lee SG, Ring D. Online health information seeking in hand and upper extremity surgery. J Hand Surg Am 2016;41(12):e469-e475.
- 2. Daraz L, Morrow AS, Ponce OJ, et al. Readability of online health information: a meta-narrative systematic review. Am J Med Qual 2018;33(5):487-492.
- 3. Aydin S, Karabacak M, Vlachos V, Margetis K. Large language models in patient education: a scoping review of applications in medicine. Front Med (Lausanne) 2024;11:1477898.
- 4. Jagiella-Lodise O, Suh N, Zelenski NA. Can patients rely on ChatGPT to answer hand pathology-related medical questions? Hand (N Y) 2025;20(5):801-809.
- 5. Croen BJ, Abdullah MS, Berns E, et al. Evaluation of patient education materials from large-language artificial intelligence models on carpal tunnel release. Hand (N Y) 2025;20(6):893-899.
- 6. Parmar RP, Daulat SR, Shah R, Brady TT, Montague M, Roth C. Readability, accuracy, and lexical diversity of new ChatGPT models for common carpal tunnel syndrome questions. Hand (N Y). Published online February 8, 2026.
- 7. Pohl NB, Derector E, Rivlin M, et al. A quality and readability comparison of artificial intelligence and popular health website education materials for common hand surgery procedures. Hand Surg Rehabil 2024;43(3):101723.
- 8. Gorgos P, Ternell KH, Hammarstrand C, et al. ChatGPT and Claude in hand surgery: an explanatory evaluation of clinical decision support on common surgical cases. Hand Surg Rehabil 2025;44(6):102530.
Ayrıntılar
Birincil Dil
İngilizce
Konular
El Cerrahisi
Bölüm
Araştırma Makalesi
Yazarlar
Hakan Ertem
Bu kişi benim
0000-0002-1724-4138
Türkiye
Yayımlanma Tarihi
22 Eylül 2026
Gönderilme Tarihi
28 Temmuz 2026
Kabul Tarihi
14 Ağustos 2026
Yayımlandığı Sayı
Yıl 2026 Cilt: 48 Sayı: 6