Research Article

Evaluating Large Language Models for Biomedical Text Summarization: A Study of Cardiovascular Research

Volume: 9 Number: 4 September 30, 2026
EN TR

Evaluating Large Language Models for Biomedical Text Summarization: A Study of Cardiovascular Research

Abstract

In this study, presented a comprehensive evaluation of abstractive and extractive summarization performance across three prominent large language models (LLMs): ChatGPT, DeepSeek, and Gemini. A total of 8,000 cardiovascular-related research abstracts were collected from PubMed and summarized using two distinct prompting strategies: abstractive and extractive. This process yielded a dataset of 48,000 summaries. To assess summarization quality, applied a multi-metric evaluation framework including semantic similarity (SBERT cosine), BLEU, GLEU, ROUGE-1 F1, ROUGE-2 F1, ROUGE-L F1, and METEOR. The results indicate that extractive summaries, particularly those generated by ChatGPT, consistently achieve higher scores across most metrics, suggesting stronger lexical fidelity and sequence retention. While Gemini shows balanced performance between abstraction and extraction, DeepSeek yields lower scores in both approaches. This work highlights critical differences in LLM behavior depending on the summarization method and offers a benchmark dataset and evaluation pipeline for future research on AI-assisted biomedical summarization.

Keywords

Supporting Institution

NA

Project Number

NA

Ethical Statement

NA

Thanks

NA

References

  1. U.S. National Library of Medicine, "Search results for cardiovascular diseases," [Online]. Available: https://pubmed.ncbi.nlm.nih.gov/?term=cardiovascular+diseases. [Accessed: 27-Jun-2025].
  2. Z. Lu, "PubMed and beyond: a survey of web tools for searching biomedical literature," Database (Oxford), 2011, Art. no. baq036, doi: 10.1093/database/baq036.
  3. B. C. Wallace, J. Kuiper, A. Sharma, M. B. Zhu, and I. J. Marshall, "Extracting PICO sentences from clinical trial reports using supervised distant supervision," Journal of Machine Learning Research, vol. 17, Art. no. 132, 2016.
  4. A. Rahman et al., "Comparative analysis based on DeepSeek, ChatGPT, and Google Gemini: Features, techniques, performance, future prospects," arXiv preprint, arXiv:2503.04783, 2025.
  5. H. Zhang, X. Liu, and J. Zhang, "Extractive summarization via ChatGPT for faithful summary generation," Findings of the Association for Computational Linguistics: EMNLP 2023, Singapore, Dec. 6–10, 2023.
  6. M. Wang, M. Wang, F. Yu, Y. Yang, J. Walker, and J. Mostafa, "A systematic review of automatic text summarization for biomedical literature and EHRs," Journal of the American Medical Informatics Association, vol. 28, no. 10, pp. 2287–2297, 2021, doi: 10.1093/jamia/ocab143.
  7. A. Chaves, C. Kesiku, and B. Garcia-Zapirain, "Automatic text summarization of biomedical text data: A systematic review," Information, vol. 13, no. 8, Art. no. 393, 2022, doi: 10.3390/info13080393.
  8. N. Kemaloğlu Alagöz and E. U. Küçüksille, "System of automatic scientific article summarization in Turkish," Pamukkale University Journal of Engineering Sciences, vol. 30, no. 4, pp. 470–481, 2024.

Details

Primary Language

English

Subjects

Computing Applications in Health

Journal Section

Research Article

Publication Date

September 30, 2026

Submission Date

September 8, 2025

Acceptance Date

April 15, 2026

Published in Issue

Year 2026 Volume: 9 Number: 4

APA
Baştürk, B., & Onan, A. (2026). Evaluating Large Language Models for Biomedical Text Summarization: A Study of Cardiovascular Research. Sakarya University Journal of Computer and Information Sciences, 9(4), 1243-1254. https://doi.org/10.35377/saucis...1780353
AMA
1.Baştürk B, Onan A. Evaluating Large Language Models for Biomedical Text Summarization: A Study of Cardiovascular Research. SAUCIS. 2026;9(4):1243-1254. doi:10.35377/saucis.1780353
Chicago
Baştürk, Burcu, and Aytuğ Onan. 2026. “Evaluating Large Language Models for Biomedical Text Summarization: A Study of Cardiovascular Research”. Sakarya University Journal of Computer and Information Sciences 9 (4): 1243-54. https://doi.org/10.35377/saucis. 1780353.
EndNote
Baştürk B, Onan A (September 1, 2026) Evaluating Large Language Models for Biomedical Text Summarization: A Study of Cardiovascular Research. Sakarya University Journal of Computer and Information Sciences 9 4 1243–1254.
IEEE
[1]B. Baştürk and A. Onan, “Evaluating Large Language Models for Biomedical Text Summarization: A Study of Cardiovascular Research”, SAUCIS, vol. 9, no. 4, pp. 1243–1254, Sept. 2026, doi: 10.35377/saucis...1780353.
ISNAD
Baştürk, Burcu - Onan, Aytuğ. “Evaluating Large Language Models for Biomedical Text Summarization: A Study of Cardiovascular Research”. Sakarya University Journal of Computer and Information Sciences 9/4 (September 1, 2026): 1243-1254. https://doi.org/10.35377/saucis. 1780353.
JAMA
1.Baştürk B, Onan A. Evaluating Large Language Models for Biomedical Text Summarization: A Study of Cardiovascular Research. SAUCIS. 2026;9:1243–1254.
MLA
Baştürk, Burcu, and Aytuğ Onan. “Evaluating Large Language Models for Biomedical Text Summarization: A Study of Cardiovascular Research”. Sakarya University Journal of Computer and Information Sciences, vol. 9, no. 4, Sept. 2026, pp. 1243-54, doi:10.35377/saucis. 1780353.
Vancouver
1.Burcu Baştürk, Aytuğ Onan. Evaluating Large Language Models for Biomedical Text Summarization: A Study of Cardiovascular Research. SAUCIS. 2026 Sep. 1;9(4):1243-54. doi:10.35377/saucis. 1780353

 

INDEXING & ABSTRACTING & ARCHIVING

 

31045 31044   Anadolu Türk Eğitim Dergisi  31047 

31043 28939 28938 34240
 

 

29070    The papers in this journal are licensed under a Creative Commons Attribution-NonCommercial 4.0 International License