Large Language Models in Medical Imaging: Increasing Use, Persistent Risks
Abstract
The use of large language models (LLMs) in medical practice, including radiological image interpretation and clinical decision support, has increased substantially in recent years. While these systems offer significant potential for improving efficiency and accessibility, important limitations remain. A key concern is the tendency of LLMs to generate plausible but incorrect information, commonly referred to as hallucination. Furthermore, such errors are often presented with high confidence, reflecting poor calibration between model certainty and accuracy. This issue is particularly critical in medical imaging, where diagnostic precision is essential. Recent studies, including our large-scale evaluation of multimodal LLMs for pneumothorax detection, demonstrate that these models may produce incorrect yet highly confident outputs. The increasing direct use of LLMs by both patients and clinicians further amplifies the risk of misinterpretation and automation bias. Given these concerns, LLM outputs should be interpreted cautiously and verified against established clinical standards. Future efforts should focus on improving model reliability, calibration, and regulatory oversight to ensure safe integration into clinical practice.
References
- 1. Thirunavukarasu AJ, Ting DSJ, Elangovan K, Gutierrez L, Tan TF, Ting DSW. Large language models in medicine. Nat Med. 2023 Aug;29(8):1930–40. doi: 10.1038/s41591-023-02448-8.
- 2. Singhal K, Azizi S, Tu T, et al. Large language models encode clinical knowledge. Nature. 2023 Aug;620(7972):172–80. doi: 10.1038/s41586-023-06291-2.
- 3. Ji Z, Lee N, Frieske R, et al. Survey of hallucination in natural language generation. ACM Comput Surv. 2023;55(12):248:1–38. doi: 10.1145/3571730.
- 4. Alkaissi H, McFarlane SI. Artificial hallucinations in ChatGPT: implications in scientific writing. Cureus. 2023 Feb 19;15(2):e35179. doi: 10.7759/cureus.35179.
- 5. Guo C, Pleiss G, Sun Y, Weinberger KQ. On Calibration of Modern Neural Networks. Proc Mach Learn Res. 2017;70:1321–30.
- 6. Minderer M, Djolonga J, Romijnders R, et al. Revisiting the Calibration of Modern Neural Networks. Adv Neural Inf Process Syst. 2021;34:15682–94.
- 7. Moor M, Banerjee O, Abad ZSH, et al. Foundation models for generalist medical artificial intelligence. Nature. 2023;616:259–65. doi: 10.1038/s41586-023-05881-4.
- 8. Acosta JN, Falcone GJ, Rajpurkar P, Topol EJ. Multimodal biomedical AI. Nat Med. 2022 Sep;28(9):1773–84. doi: 10.1038/s41591-022-01981-2.
Details
Primary Language
English
Subjects
Natural Language Processing, Diagnostic Radiography
Journal Section
Letter to Editor
Authors
Publication Date
August 29, 2026
Submission Date
April 2, 2026
Acceptance Date
May 29, 2026
Published in Issue
Year 2026 Volume: 6 Number: 2