Research Article

How Artificial, How Intelligent? LLMs and Medical Experts Face Off in TMA Diagnosis

Number: Advanced Online Publication Early Pub Date: September 2, 2026
EN

How Artificial, How Intelligent? LLMs and Medical Experts Face Off in TMA Diagnosis

Abstract

Objective: To compare physician specialists and large language models in the differential diagnosis and plasma exchange decision making of standardized thrombotic microangiopathy case vignettes.
Methods: This case vignette based comparative study included three standardized TMA scenarios representing atypical hemolytic uremic syndrome, malignant hypertension associated TMA, and acquired thrombotic thrombocytopenic purpura. Nine evaluators participated: two nephrologists, two hematologists, two internal medicine specialists, and three large language models (ChatGPT-5.2, Gemini Pro, and Microsoft Copilot). For each case, participants provided a most likely diagnosis, three differential diagnoses, supporting clinical reasoning, and a plasma exchange decision. Diagnostic accuracy, plasma exchange appropriateness, and critical management errors were analyzed descriptively and comparatively.
Results: A total of 27 independent evaluations were analyzed. Overall Top-1 diagnostic accuracy was 51.9%. Accuracy varied by case, with the lowest performance observed in the aHUS scenario (22.2%) and the highest in malignant hypertension–associated TMA (77.8%). Physicians achieved a Top-1 accuracy of 50.0% (9/18), whereas LLMs demonstrated 55.6% accuracy (5/9). Plasma exchange decision accuracy was 63.0% overall, with overtreatment more common in non-TTP cases. Ten critical management errors (37.0%) were identified. Agreement for plasma exchange decisions demonstrated fair concordance (κ = 0.29).
Conclusions: Substantial variability exists in both diagnostic classification and therapeutic decision-making in TMA scenarios. While LLMs demonstrated competitive diagnostic performance in selected cases, management variability persisted across evaluator groups. These findings highlight the diagnostic complexity of TMA and support the role of artificial intelligence as an adjunctive clinical decision-support tool rather than an autonomous decision-maker.

Keywords

Ethical Statement

This study was approved by the Scientific Research Ethics Committee of Antalya Training and Research Hospital (Decision No: 3/28; Date: February 13, 2025). The study was conducted in accordance with the principles of the Declaration of Helsinki.

References

  1. Story CM, Gerber GF, Chaturvedi S. Medical consult: aHUS, TTP? How to distinguish and what to do. Hematology Am Soc Hematol Educ Program. 2023;2023(1):745-53. doi:10.1182/hematology.2023000501.
  2. Papakonstantinou A, Kalmoukos P, Mpalaska A, Koravou EE, Gavriilaki E. ADAMTS13 in the new era of TTP. Int J Mol Sci. 2024;25(15):8137. doi:10.3390/ijms25158137.
  3. Khanal N, Dahal S, Upadhyay S, Bhatt VR, Bierman PJ. Differentiating malignant hypertension-induced thrombotic microangiopathy from thrombotic thrombocytopenic purpura. Ther Adv Hematol. 2015;6(3):97-102. doi:10.1177/2040620715571076.
  4. Addad VV, Palma LMP, Vaisbich MH, Pacheco Barbosa AM, da Rocha NC, de Almeida Cardoso MM, et al. A comprehensive model for assessing and classifying patients with thrombotic microangiopathy: the TMA-INSIGHT score. Thromb J. 2023;21:119. doi:10.1186/s12959-023-00564-6.
  5. Abou-Ismail MY, Zhang C, Presson AP, Chaturvedi S, Antun AG, Farland AM, et al. A machine learning approach to predict mortality due to immune-mediated thrombotic thrombocytopenic purpura. Res Pract Thromb Haemost. 2024;8(3):102388. doi:10.1016/j.rpth.2024.102388.
  6. Herrera Rivera N, McClintock DS, Alterman MA, Alterman TAL, Pruitt HD, Olsen GM, et al. A clinical laboratorian's journey in developing a machine learning algorithm to assist in testing utilization and stewardship. J Lab Precis Med. 2023;8:22. doi:10.21037/jlpm-23-11.
  7. Zhang Z, Yang S, Wang X. Schistocyte detection in artificial intelligence age. Int J Lab Hematol. 2024;46(3):427-33. doi:10.1111/ijlh.14260.
  8. Hirosawa T, Kawamura R, Harada Y, Mizuta K, Tokumasu K, Kaji Y, et al. ChatGPT-generated differential diagnosis lists for complex case-derived clinical vignettes: diagnostic accuracy evaluation. JMIR Med Inform. 2023;11:e48808. doi:10.2196/48808.

Details

Primary Language

English

Subjects

​Internal Diseases

Journal Section

Research Article

Early Pub Date

September 2, 2026

Publication Date

-

Submission Date

February 15, 2026

Acceptance Date

May 1, 2026

Published in Issue

Year 2026 Number: Advanced Online Publication

APA
Zorlu Görgülügil, G., Kaçar, C., Şen, E., İnci, A., & Karakuş, V. (2026). How Artificial, How Intelligent? LLMs and Medical Experts Face Off in TMA Diagnosis. Sakarya Medical Journal, Advanced Online Publication. https://doi.org/10.31832/smj.1889486
AMA
1.Zorlu Görgülügil G, Kaçar C, Şen E, İnci A, Karakuş V. How Artificial, How Intelligent? LLMs and Medical Experts Face Off in TMA Diagnosis. Sakarya Medical Journal. 2026;(Advanced Online Publication). doi:10.31832/smj.1889486
Chicago
Zorlu Görgülügil, Gizem, Ceyda Kaçar, Elif Şen, Ayça İnci, and Volkan Karakuş. 2026. “How Artificial, How Intelligent? LLMs and Medical Experts Face Off in TMA Diagnosis”. Sakarya Medical Journal, no. Advanced Online Publication. https://doi.org/10.31832/smj.1889486.
EndNote
Zorlu Görgülügil G, Kaçar C, Şen E, İnci A, Karakuş V (September 1, 2026) How Artificial, How Intelligent? LLMs and Medical Experts Face Off in TMA Diagnosis. Sakarya Medical Journal Advanced Online Publication
IEEE
[1]G. Zorlu Görgülügil, C. Kaçar, E. Şen, A. İnci, and V. Karakuş, “How Artificial, How Intelligent? LLMs and Medical Experts Face Off in TMA Diagnosis”, Sakarya Medical Journal, no. Advanced Online Publication, Sept. 2026, doi: 10.31832/smj.1889486.
ISNAD
Zorlu Görgülügil, Gizem - Kaçar, Ceyda - Şen, Elif - İnci, Ayça - Karakuş, Volkan. “How Artificial, How Intelligent? LLMs and Medical Experts Face Off in TMA Diagnosis”. Sakarya Medical Journal. Advanced Online Publication (September 1, 2026). https://doi.org/10.31832/smj.1889486.
JAMA
1.Zorlu Görgülügil G, Kaçar C, Şen E, İnci A, Karakuş V. How Artificial, How Intelligent? LLMs and Medical Experts Face Off in TMA Diagnosis. Sakarya Medical Journal. 2026. doi:10.31832/smj.1889486.
MLA
Zorlu Görgülügil, Gizem, et al. “How Artificial, How Intelligent? LLMs and Medical Experts Face Off in TMA Diagnosis”. Sakarya Medical Journal, no. Advanced Online Publication, Sept. 2026, doi:10.31832/smj.1889486.
Vancouver
1.Gizem Zorlu Görgülügil, Ceyda Kaçar, Elif Şen, Ayça İnci, Volkan Karakuş. How Artificial, How Intelligent? LLMs and Medical Experts Face Off in TMA Diagnosis. Sakarya Medical Journal. 2026 Sep. 1;(Advanced Online Publication). doi:10.31832/smj.1889486

INDEXING & ABSTRACTING & ARCHIVING

 

download?token=eyJhdXRoX3JvbGVzIjpbXSwiZW5kcG9pbnQiOiJqb3VybmFsIiwib3JpZ2luYWxuYW1lIjoiU2NvcHVzLnBuZyIsInBhdGgiOiI5NzZkLzU5ZjkvMjAyMC82YTkxMmNmOGRiMTQzMS44MDUwMjExMS5wbmciLCJleHAiOjE3ODc5MDI3MjksIm5vbmNlIjoiMGM1Y2FmMmZlZjRlYWIyMTEyMGVjYjUyZTU2NWEyMzAifQ.c7PzOMe8mvOEUJa4kmFHwDEXwzC93EDJLv4c7YpMba0  29985  30950  30951 30954 34273
 

 

30703 The published articles in SMJ are licensed under a Creative Commons Attribution-NonCommercial 4.0 International License.