Araştırma Makalesi

Comparative Performance Evaluation of Multimodal Large Language Models, Radiologist, and Anatomist in Visual Neuroanatomy Questions

Cilt: 50 Sayı: 3 12 Ocak 2025
PDF İndir
TR EN

Comparative Performance Evaluation of Multimodal Large Language Models, Radiologist, and Anatomist in Visual Neuroanatomy Questions

Öz

This study examined the performance of four different multimodal Large Language Models (LLMs)—GPT4-V, GPT-4o, LLaVA, and Gemini 1.5 Flash—on multiple-choice visual neuroanatomy questions, comparing them to a radiologist and an anatomist. The study employed a cross-sectional design and evaluated responses to 100 visual questions sourced from the Radiopaedia website. The accuracy of the responses was analyzed using the McNemar test. According to the results, the radiologist demonstrated the highest performance with an accuracy rate of 90%, while the anatomist achieved an accuracy rate of 67%. Among the multimodal LLMs, GPT-4o performed the best, with an accuracy rate of 45%, followed by Gemini 1.5 Flash at 35%, ChatGPT4-V at 22%, and LLaVA at 15%. The radiologist significantly outperformed both the anatomist and all multimodal LLMs (p<0.001). GPT-4o significantly outperformed GPT4-V and LLaVA (p<0.001), but no significant difference was found between GPT-4o and Gemini 1.5 Flash (p=0.123). However, Gemini 1.5 Flash showed significant superiority over LLaVA (p<0.001) and also demonstrated a statistically significant difference compared to GPT4-V (p=0.004). This study highlights the significant performance gap between multimodal LLMs and medical professionals. While multimodal LLMs hold great potential in the medical field, they have not yet reached the level of accuracy of medical experts in correctly identifying neuroanatomical regions.

Anahtar Kelimeler

Etik Beyan

Ethics Committee Approval Information This study does not require ethics committee approval as it was conducted using publicly available internet data, and the images do not contain patient information. The study was carried out in accordance with the Standards for Reporting of Diagnostic Accuracy (STARD) and the Checklist for Artificial Intelligence in Medical Imaging (CLAIM).

Teşekkür

The authors used ChatGPT 4o (September 2024 Release; OpenAI; https://chat.openai.com/) to review grammar and English translation. The content of the publication is entirely the responsibility of the authors, who have reviewed and edited it as they deemed necessary. The authors thank Juliette Hancox, Visual Licensing Manager at Radiopaedia.org, for granting permission to use the images from the associated website.

Kaynakça

  1. 1. Clusmann J, Kolbinger FR, Muti HS, Carrero ZI, Eckardt JN, Laleh NG, Löffler CML, Schwarzkopf SC, Unger M, Veldhuizen GP, Wagner SJ, Kather JN (2023) The future landscape of large language models in medicine. Commun Med (Lond) 3:141. https://doi.org/10.1038/s43856-023-00370-1
  2. 2. GPT-4 is OpenAI’s most advanced system, producing safer and more useful responses. OpenAI.https://openai.com/gpt-4 / GPT-4V(ision) System Card. OpenAI. Accessed Date Accessed
  3. 3. Liu H, Li C, Wu Q, Lee YJ (2024) Visual instruction tuning. Adv Neural Inf Process Syst 36
  4. 4. https://deepmind.google/technologies/gemini/flash/. Accessed Date Accessed
  5. 5. https://openai.com/index/hello-gpt-4o/. Accessed Date Accessed
  6. 6. Kuang Y-R, Zou M-X, Niu H-Q, Zheng B-Y, Zhang T-L, Zheng B-W (2023) ChatGPT encounters multiple opportunities and challenges in neurosurgery. Int J Surg 109:2886-2891. https://doi.org/doi: 10.1097/JS9.0000000000000571
  7. 7. Gunes YC, Camur E, Cesur T (2024) Correspondence on ‘Evaluation of ChatGPT in knowledge of newly evolving neurosurgery: middle meningeal artery embolization for subdural hematoma management’by Koester et al. J Neurointerv Surg
  8. 8. Andykarayalar R, Surapaneni KM (2024) ChatGPT in Pediatrics: Unraveling Its Significance as a Clinical Decision Support Tool. Indian Pediatr 61:357-358

Ayrıntılar

Birincil Dil

İngilizce

Konular

Radyoloji ve Organ Görüntüleme, Anatomi

Bölüm

Araştırma Makalesi

Yayımlanma Tarihi

12 Ocak 2025

Gönderilme Tarihi

16 Ekim 2024

Kabul Tarihi

2 Ocak 2025

Yayımlandığı Sayı

Yıl 2024 Cilt: 50 Sayı: 3

Kaynak Göster

APA
Güneş, Y. C., & Ülkir, M. (2025). Comparative Performance Evaluation of Multimodal Large Language Models, Radiologist, and Anatomist in Visual Neuroanatomy Questions. Journal of Uludağ University Medical Faculty, 50(3), 551-556. https://doi.org/10.32708/uutfd.1568479
AMA
1.Güneş YC, Ülkir M. Comparative Performance Evaluation of Multimodal Large Language Models, Radiologist, and Anatomist in Visual Neuroanatomy Questions. Uludağ Tıp Derg. 2025;50(3):551-556. doi:10.32708/uutfd.1568479
Chicago
Güneş, Yasin Celal, ve Mehmet Ülkir. 2025. “Comparative Performance Evaluation of Multimodal Large Language Models, Radiologist, and Anatomist in Visual Neuroanatomy Questions”. Journal of Uludağ University Medical Faculty 50 (3): 551-56. https://doi.org/10.32708/uutfd.1568479.
EndNote
Güneş YC, Ülkir M (01 Ocak 2025) Comparative Performance Evaluation of Multimodal Large Language Models, Radiologist, and Anatomist in Visual Neuroanatomy Questions. Journal of Uludağ University Medical Faculty 50 3 551–556.
IEEE
[1]Y. C. Güneş ve M. Ülkir, “Comparative Performance Evaluation of Multimodal Large Language Models, Radiologist, and Anatomist in Visual Neuroanatomy Questions”, Uludağ Tıp Derg, c. 50, sy 3, ss. 551–556, Oca. 2025, doi: 10.32708/uutfd.1568479.
ISNAD
Güneş, Yasin Celal - Ülkir, Mehmet. “Comparative Performance Evaluation of Multimodal Large Language Models, Radiologist, and Anatomist in Visual Neuroanatomy Questions”. Journal of Uludağ University Medical Faculty 50/3 (01 Ocak 2025): 551-556. https://doi.org/10.32708/uutfd.1568479.
JAMA
1.Güneş YC, Ülkir M. Comparative Performance Evaluation of Multimodal Large Language Models, Radiologist, and Anatomist in Visual Neuroanatomy Questions. Uludağ Tıp Derg. 2025;50:551–556.
MLA
Güneş, Yasin Celal, ve Mehmet Ülkir. “Comparative Performance Evaluation of Multimodal Large Language Models, Radiologist, and Anatomist in Visual Neuroanatomy Questions”. Journal of Uludağ University Medical Faculty, c. 50, sy 3, Ocak 2025, ss. 551-6, doi:10.32708/uutfd.1568479.
Vancouver
1.Yasin Celal Güneş, Mehmet Ülkir. Comparative Performance Evaluation of Multimodal Large Language Models, Radiologist, and Anatomist in Visual Neuroanatomy Questions. Uludağ Tıp Derg. 01 Ocak 2025;50(3):551-6. doi:10.32708/uutfd.1568479

Cited By

ISSN: 1300-414X, e-ISSN: 2645-9027

Uludağ Üniversitesi Tıp Fakültesi Dergisi "Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License" ile lisanslanmaktadır.


Creative Commons License
Journal of Uludag University Medical Faculty is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License.

2023