EN
TR
Evaluating Large Language Models for Biomedical Text Summarization: A Study of Cardiovascular Research
Abstract
In this study, presented a comprehensive evaluation of abstractive and extractive summarization performance across three prominent large language models (LLMs): ChatGPT, DeepSeek, and Gemini. A total of 8,000 cardiovascular-related research abstracts were collected from PubMed and summarized using two distinct prompting strategies: abstractive and extractive. This process yielded a dataset of 48,000 summaries. To assess summarization quality, applied a multi-metric evaluation framework including semantic similarity (SBERT cosine), BLEU, GLEU, ROUGE-1 F1, ROUGE-2 F1, ROUGE-L F1, and METEOR. The results indicate that extractive summaries, particularly those generated by ChatGPT, consistently achieve higher scores across most metrics, suggesting stronger lexical fidelity and sequence retention. While Gemini shows balanced performance between abstraction and extraction, DeepSeek yields lower scores in both approaches. This work highlights critical differences in LLM behavior depending on the summarization method and offers a benchmark dataset and evaluation pipeline for future research on AI-assisted biomedical summarization.
Keywords
- Large Language Models (LLMs)
- Abstractive Summarization
- Extractive Summarization
- Cardiovascular Research
- Natural Language Processing(NLP)
- Evaluation Metrics
Supporting Institution
NA
Project Number
NA
Ethical Statement
NA
Thanks
NA
References
- U.S. National Library of Medicine, "Search results for cardiovascular diseases," [Online]. Available: https://pubmed.ncbi.nlm.nih.gov/?term=cardiovascular+diseases. [Accessed: 27-Jun-2025].
- Z. Lu, "PubMed and beyond: a survey of web tools for searching biomedical literature," Database (Oxford), 2011, Art. no. baq036, doi: 10.1093/database/baq036.
- B. C. Wallace, J. Kuiper, A. Sharma, M. B. Zhu, and I. J. Marshall, "Extracting PICO sentences from clinical trial reports using supervised distant supervision," Journal of Machine Learning Research, vol. 17, Art. no. 132, 2016.
- A. Rahman et al., "Comparative analysis based on DeepSeek, ChatGPT, and Google Gemini: Features, techniques, performance, future prospects," arXiv preprint, arXiv:2503.04783, 2025.
- H. Zhang, X. Liu, and J. Zhang, "Extractive summarization via ChatGPT for faithful summary generation," Findings of the Association for Computational Linguistics: EMNLP 2023, Singapore, Dec. 6–10, 2023.
- M. Wang, M. Wang, F. Yu, Y. Yang, J. Walker, and J. Mostafa, "A systematic review of automatic text summarization for biomedical literature and EHRs," Journal of the American Medical Informatics Association, vol. 28, no. 10, pp. 2287–2297, 2021, doi: 10.1093/jamia/ocab143.
- A. Chaves, C. Kesiku, and B. Garcia-Zapirain, "Automatic text summarization of biomedical text data: A systematic review," Information, vol. 13, no. 8, Art. no. 393, 2022, doi: 10.3390/info13080393.
- N. Kemaloğlu Alagöz and E. U. Küçüksille, "System of automatic scientific article summarization in Turkish," Pamukkale University Journal of Engineering Sciences, vol. 30, no. 4, pp. 470–481, 2024.
Details
Primary Language
English
Subjects
Computing Applications in Health
Journal Section
Research Article
Publication Date
September 30, 2026
Submission Date
September 8, 2025
Acceptance Date
April 15, 2026
Published in Issue
Year 2026 Volume: 9 Number: 4
APA
Baştürk, B., & Onan, A. (2026). Evaluating Large Language Models for Biomedical Text Summarization: A Study of Cardiovascular Research. Sakarya University Journal of Computer and Information Sciences, 9(4), 1243-1254. https://doi.org/10.35377/saucis...1780353
AMA
1.Baştürk B, Onan A. Evaluating Large Language Models for Biomedical Text Summarization: A Study of Cardiovascular Research. SAUCIS. 2026;9(4):1243-1254. doi:10.35377/saucis.1780353
Chicago
Baştürk, Burcu, and Aytuğ Onan. 2026. “Evaluating Large Language Models for Biomedical Text Summarization: A Study of Cardiovascular Research”. Sakarya University Journal of Computer and Information Sciences 9 (4): 1243-54. https://doi.org/10.35377/saucis. 1780353.
EndNote
Baştürk B, Onan A (September 1, 2026) Evaluating Large Language Models for Biomedical Text Summarization: A Study of Cardiovascular Research. Sakarya University Journal of Computer and Information Sciences 9 4 1243–1254.
IEEE
[1]B. Baştürk and A. Onan, “Evaluating Large Language Models for Biomedical Text Summarization: A Study of Cardiovascular Research”, SAUCIS, vol. 9, no. 4, pp. 1243–1254, Sept. 2026, doi: 10.35377/saucis...1780353.
ISNAD
Baştürk, Burcu - Onan, Aytuğ. “Evaluating Large Language Models for Biomedical Text Summarization: A Study of Cardiovascular Research”. Sakarya University Journal of Computer and Information Sciences 9/4 (September 1, 2026): 1243-1254. https://doi.org/10.35377/saucis. 1780353.
JAMA
1.Baştürk B, Onan A. Evaluating Large Language Models for Biomedical Text Summarization: A Study of Cardiovascular Research. SAUCIS. 2026;9:1243–1254.
MLA
Baştürk, Burcu, and Aytuğ Onan. “Evaluating Large Language Models for Biomedical Text Summarization: A Study of Cardiovascular Research”. Sakarya University Journal of Computer and Information Sciences, vol. 9, no. 4, Sept. 2026, pp. 1243-54, doi:10.35377/saucis. 1780353.
Vancouver
1.Burcu Baştürk, Aytuğ Onan. Evaluating Large Language Models for Biomedical Text Summarization: A Study of Cardiovascular Research. SAUCIS. 2026 Sep. 1;9(4):1243-54. doi:10.35377/saucis. 1780353