Research Article

Evaluation of Large Language Models for Radiology Report Conclusion Generation

Volume: 6 Number: 2 August 29, 2026
TR EN

Evaluation of Large Language Models for Radiology Report Conclusion Generation

Abstract

Purpose: To evaluate and compare the performance of ChatGPT 5.5 and Gemini 3.5 Flash in automatically generating conclusion sections from Turkish non-contrast thorax computed tomography (CT) findings using natural language processing metrics. Materials and Methods: In this retrospective study, "Findings" sections from 50 consecutive radiologist-authored non-contrast thorax CT reports were processed by ChatGPT and Gemini via API. A standardized, rule-based prompt was utilized. The LLM-generated conclusions were quantitatively compared to the original expert radiologist conclusions (reference standard) using a Turkish-adapted BERTScore for semantic similarity, ROUGE-L for structural overlap, and a text length ratio to assess clinical conciseness. Results: ChatGPT demonstrated a significantly higher overall BERTScore F1 metric (0.6780 ± 0.0712 vs. 0.6427 ± 0.0881, p < 0.001) and precision (0.6178 ± 0.0757 vs. 0.5612 ± 0.0864, p < 0.001) compared to Gemini. Recall was comparable between the two models (p = 0.9719). Structural similarity via ROUGE-L F1 showed no significant difference (p = 0.5656). Both models generated conclusion sections significantly longer than the original human reports (ratio > 1.0); however, the text length ratio of ChatGPT was significantly lower, indicating a more concise output compared to Gemini (2.4287 vs. 3.0564, p < 0.001). Conclusion: Both models demonstrate the capacity to synthesize clinical summaries, with ChatGPT offering superior semantic precision for Turkish radiology texts. However, the verbosity of both models compared to the compact reporting style of human radiologists, and comparably lower performance on semantic and structural scores necessitates a human-in-the-loop approach for safe and efficient integration into clinical workflows.

Keywords

Supporting Institution

N/A

Ethical Statement

Ethical approval for this study was obtained from the İzmir Kâtip Çelebi Health Research Institutional Review Board on June 11, 2026 (IRB #: 0513). The institutional review board waived the requirement for informed consent due to the retrospective nature of the study and the use of non-identifiable data.

Thanks

N/A

References

  1. 1. Flanders AE, Lakhani P. Radiology reporting and communications: a look forward. Neuroimaging Clin N Am. 2012;22:477-96. doi: 10.1016/j.nic.2012.04.009.
  2. 2. Bosmans JML, Weyler JJ, De Schepper AM, Parizel PM. The radiology report as seen by radiologists and referring clinicians: results of the COVER and ROVER surveys. Radiology. 2011;259:184-95. doi: 10.1148/radiol.10101045.
  3. 3. Reiner BI, Knight N, Siegel EL. Radiology reporting, past, present, and future: the radiologist's perspective. J Am Coll Radiol. 2007;4:313-9. doi: 10.1016/j.jacr.2007.01.015.
  4. 4. Hartung MP, Bickle IC, Gaillard F, Kanne JP. How to create a great radiology report. Radiographics. 2020;40:1658-70. doi: 10.1148/rg.2020200020.
  5. 5. Lukaszewicz A, Uricchio J, Gerasymchuk G. The art of the radiology report: practical and stylistic guidelines for perfecting the conveyance of imaging findings. Can Assoc Radiol J. 2016;67:318-21. doi: 10.1016/j.carj.2016.03.001.
  6. 6. Salbas A, Kul Baysan E. Assessment of large language models in musculoskeletal radiological anatomy: a comparative study with radiologists. Jt Dis Relat Surg. 2026;37:190-9. doi: 10.52312/jdrs.2026.2436.
  7. 7. Güzel HE, Oleaga L, Koç AM, Junquero V, Merino C. Large language models solving the European Diploma in Radiology: a comparative evaluation. Acad Radiol. 2026;33:1871-8. doi: 10.1016/j.acra.2026.01.040
  8. 8. Büyüktoka RE, Surucu M, Erekli Derinkaya PB, et al. Applying large language model for automated quality scoring of radiology requisitions using a standardized criteria. Eur Radiol. Epub ahead of print 2025. doi: 10.1007/s00330-025-11933-2.

Details

Primary Language

English

Subjects

Radiology and Organ Imaging

Journal Section

Research Article

Publication Date

August 29, 2026

Submission Date

July 9, 2026

Acceptance Date

August 24, 2026

Published in Issue

Year 2026 Volume: 6 Number: 2

APA
Büyüktoka, R. E., & Büyüktoka, A. D. (2026). Evaluation of Large Language Models for Radiology Report Conclusion Generation. Sağlık Bilimlerinde Yapay Zeka Dergisi, 6(2), 25-31. https://doi.org/10.52309/jaihs.1991112
AMA
1.Büyüktoka RE, Büyüktoka AD. Evaluation of Large Language Models for Radiology Report Conclusion Generation. JAIHS. 2026;6(2):25-31. doi:10.52309/jaihs.1991112
Chicago
Büyüktoka, Raşit Eren, and Aslı Dilara Büyüktoka. 2026. “Evaluation of Large Language Models for Radiology Report Conclusion Generation”. Sağlık Bilimlerinde Yapay Zeka Dergisi 6 (2): 25-31. https://doi.org/10.52309/jaihs.1991112.
EndNote
Büyüktoka RE, Büyüktoka AD (August 1, 2026) Evaluation of Large Language Models for Radiology Report Conclusion Generation. Sağlık Bilimlerinde Yapay Zeka Dergisi 6 2 25–31.
IEEE
[1]R. E. Büyüktoka and A. D. Büyüktoka, “Evaluation of Large Language Models for Radiology Report Conclusion Generation”, JAIHS, vol. 6, no. 2, pp. 25–31, Aug. 2026, doi: 10.52309/jaihs.1991112.
ISNAD
Büyüktoka, Raşit Eren - Büyüktoka, Aslı Dilara. “Evaluation of Large Language Models for Radiology Report Conclusion Generation”. Sağlık Bilimlerinde Yapay Zeka Dergisi 6/2 (August 1, 2026): 25-31. https://doi.org/10.52309/jaihs.1991112.
JAMA
1.Büyüktoka RE, Büyüktoka AD. Evaluation of Large Language Models for Radiology Report Conclusion Generation. JAIHS. 2026;6:25–31.
MLA
Büyüktoka, Raşit Eren, and Aslı Dilara Büyüktoka. “Evaluation of Large Language Models for Radiology Report Conclusion Generation”. Sağlık Bilimlerinde Yapay Zeka Dergisi, vol. 6, no. 2, Aug. 2026, pp. 25-31, doi:10.52309/jaihs.1991112.
Vancouver
1.Raşit Eren Büyüktoka, Aslı Dilara Büyüktoka. Evaluation of Large Language Models for Radiology Report Conclusion Generation. JAIHS. 2026 Aug. 1;6(2):25-31. doi:10.52309/jaihs.1991112