Comparative Analysis of Vision Transformers and U-Net for Medical Image Segmentation in Early Disease Detection: A Deep Learning Approach
Abstract
Medical image segmentation remains a critical challenge in computer-aided diagnosis systems, particularly for early disease detection where precise boundary delineation can significantly impact patient outcomes. This study presents a comprehensive comparative analysis between Vision Transformer (ViT) based architectures and the conventional U-Net model for multi-organ segmentation tasks using chest CT scans and retinal fundus images. We evaluated both architectures on three distinct datasets comprising 15,420 annotated medical images, focusing on lung nodule detection, liver lesion segmentation, and retinal vessel segmentation for diabetic retinopathy screening. Our experimental results demonstrate that while U-Net achieves superior performance on smaller datasets (Dice coefficient: 0.89 ± 0.03), Vision Transformers exhibit remarkable capabilities with larger training samples (Dice coefficient: 0.93 ± 0.02), showing 4.5% improvement in segmentation accuracy. The ViT-based approach demonstrated enhanced generalization capabilities across diverse imaging modalities, reducing false positive rates by 31% compared to U-Net in cross-dataset validation. Furthermore, computational efficiency analysis revealed that despite requiring 2.3× more training time, ViT models reduced inference time by 18% in clinical deployment scenarios. Performance evaluation across image quality levels showed ViT maintained more consistent performance across signal-to-noise ratios (Dice drop: 4.2% from high to low SNR) compared to U-Net (8.7% drop), demonstrating transformers' robustness to image degradation in clinical settings where scan quality varies. These findings suggest that the choice between architectures should be guided by dataset size, computational resources, and specific clinical requirements, with hybrid approaches showing promising potential for future development.
Keywords
References
- Armato III, S. G., McLennan, G., Bidaut, L., et al. (2011). The Lung Image Database Consortium (LIDC) and Image Database Resource Initiative (IDRI): A completed reference database of lung nodules on CT scans. Medical Physics, 38(2), 915-931.
- Aydın, V. A. (2024). Comparison of CNN-based methods for yoga pose classification. Turkish Journal of Engineering, 8(1), 65-75. https://doi.org/10.31127/tuje.1348210
- Azad, R., Aghdam, E. K., Rauland, A., et al. (2022). Medical image segmentation review: The success of U-Net. arXiv preprint arXiv:2211.14830.
- Bai, W., Sinclair, M., Tarroni, G., et al. (2018). Automated cardiovascular magnetic resonance image analysis with fully convolutional networks. Journal of Cardiovascular Magnetic Resonance, 20(1), 65.
- Bilic, P., Christ, P. F., Vorontsov, E., et al. (2019). The Liver Tumor Segmentation Benchmark (LiTS). arXiv preprint arXiv:1901.04056.
- Cao, H., Wang, Y., Chen, J., et al. (2022). Swin-Unet: Unet-like pure transformer for medical image segmentation. European Conference on Computer Vision, 205-218.
- Chen, J., Lu, Y., Yu, Q., Luo, X., Adeli, E., Wang, Y., Lu, L., & Zhou, Y. (2021). TransUNet: Transformers make strong encoders for medical image segmentation. arXiv preprint arXiv:2102.04306. https://doi.org/10.48550/arXiv.2102.04306
- Çiçek, Ö., Abdulkadir, A., Lienkamp, S. S., Brox, T., & Ronneberger, O. (2016). 3D U-Net: Learning dense volumetric segmentation from sparse annotation. Medical Image Computing and Computer-Assisted Intervention, 424-432.
Details
Primary Language
English
Subjects
Wireless Communication Systems and Technologies (Incl. Microwave and Millimetrewave)
Journal Section
Research Article
Authors
Publication Date
May 1, 2026
Submission Date
November 13, 2025
Acceptance Date
January 4, 2026
Published in Issue
Year 2026 Volume: 10 Number: 2
APA
S.p., S., & Muthukumarasamy, C. (2026). Comparative Analysis of Vision Transformers and U-Net for Medical Image Segmentation in Early Disease Detection: A Deep Learning Approach. Turkish Journal of Engineering, 10(2), 378-395. https://doi.org/10.31127/tuje.1822987
AMA
1.S.p. S, Muthukumarasamy C. Comparative Analysis of Vision Transformers and U-Net for Medical Image Segmentation in Early Disease Detection: A Deep Learning Approach. TUJE. 2026;10(2):378-395. doi:10.31127/tuje.1822987
Chicago
S.p., Senthilkumar, and Chandramouleeswaran Muthukumarasamy. 2026. “Comparative Analysis of Vision Transformers and U-Net for Medical Image Segmentation in Early Disease Detection: A Deep Learning Approach”. Turkish Journal of Engineering 10 (2): 378-95. https://doi.org/10.31127/tuje.1822987.
EndNote
S.p. S, Muthukumarasamy C (May 1, 2026) Comparative Analysis of Vision Transformers and U-Net for Medical Image Segmentation in Early Disease Detection: A Deep Learning Approach. Turkish Journal of Engineering 10 2 378–395.
IEEE
[1]S. S.p. and C. Muthukumarasamy, “Comparative Analysis of Vision Transformers and U-Net for Medical Image Segmentation in Early Disease Detection: A Deep Learning Approach”, TUJE, vol. 10, no. 2, pp. 378–395, May 2026, doi: 10.31127/tuje.1822987.
ISNAD
S.p., Senthilkumar - Muthukumarasamy, Chandramouleeswaran. “Comparative Analysis of Vision Transformers and U-Net for Medical Image Segmentation in Early Disease Detection: A Deep Learning Approach”. Turkish Journal of Engineering 10/2 (May 1, 2026): 378-395. https://doi.org/10.31127/tuje.1822987.
JAMA
1.S.p. S, Muthukumarasamy C. Comparative Analysis of Vision Transformers and U-Net for Medical Image Segmentation in Early Disease Detection: A Deep Learning Approach. TUJE. 2026;10:378–395.
MLA
S.p., Senthilkumar, and Chandramouleeswaran Muthukumarasamy. “Comparative Analysis of Vision Transformers and U-Net for Medical Image Segmentation in Early Disease Detection: A Deep Learning Approach”. Turkish Journal of Engineering, vol. 10, no. 2, May 2026, pp. 378-95, doi:10.31127/tuje.1822987.
Vancouver
1.Senthilkumar S.p., Chandramouleeswaran Muthukumarasamy. Comparative Analysis of Vision Transformers and U-Net for Medical Image Segmentation in Early Disease Detection: A Deep Learning Approach. TUJE. 2026 May 1;10(2):378-95. doi:10.31127/tuje.1822987