Building Segmentation from High-Resolution Remote Sensing Imagery: A Comparative Evaluation of Convolutional and Transformer Architectures
Öz
Architectures for building extraction are usually compared under different training and inference set-ups, which makes the reported performance figures incommensurable. This study evaluates three convolutional segmentation architectures (U-Net, U-Net++, DeepLabV3+) and one transformer-based architecture (SegFormer) on the aerial orthophoto subset of the WHU Building Dataset under a single training and inference protocol. Fixing the encoder to ResNet-34 in the convolutional models separates the effect of decoder design from backbone capacity. U-Net++ achieved the highest performance at 89.04% IoU and 94.20% F1, and the difference between the models remained within a narrow range of 2.42 points on the standard IoU metric. On the boundary IoU metric, which measures overlap only within the edge band, the same spread rose to 9.26 points, roughly 3.8 times as wide. Reporting standard IoU alone therefore understates the difference between architectures systematically. An inference protocol combining an exponential moving average, eight-way test-time augmentation and a decision threshold tuned on the validation set yielded gains of 0.47 to 0.69 IoU points without changing the ranking of the models; that magnitude corresponds to roughly a quarter of the total spread attributable to architecture. On the cost side, parameter count proved a misleading indicator: a 6.7% difference in parameters between U-Net++ and U-Net corresponded to a 135.2% difference in multiply-accumulate operations. The results show that model selection in building extraction cannot rest on a single accuracy figure, and that boundary-based metrics should govern the choice in applications where the position of the boundary is critical.
Anahtar Kelimeler
Kaynakça
- [1] Ji S, Wei S, Lu M. Fully Convolutional Networks for Multisource Building Extraction From an Open Aerial and Satellite Imagery Data Set. IEEE Transactions on Geoscience and Remote Sensing, 57(574-586), (2019).
- [2] Sirmacek B, Unsalan C. Building detection from aerial images using invariant color features and shadow information. 2008 23rd International Symposium on Computer and Information Sciences, 1-5, (2008).
- [3] Liow Y-T, Pavlidis T. Use of shadows for extracting buildings in aerial images. Computer Vision, Graphics, and Image Processing, 49(242-277), (1990).
- [4] Long J, Shelhamer E, Darrell T. Fully Convolutional Networks for Semantic Segmentation. 3431-3440, (2015).
- [5] Ronneberger O, Fischer P, Brox T. U-Net: Convolutional Networks for Biomedical Image Segmentation. Medical Image Computing and Computer-Assisted Intervention – MICCAI 2015, Springer International Publishing, Cham, 234-241, (2015).
- [6] Zhou Z, Rahman Siddiquee MM, Tajbakhsh N, Liang J. UNet++: A Nested U-Net Architecture for Medical Image Segmentation. Deep Learning in Medical Image Analysis and Multimodal Learning for Clinical Decision Support, Springer International Publishing, Cham, 3-11, (2018).
- [7] Chen L-C, Zhu Y, Papandreou G, Schroff F, Adam H. Encoder-Decoder with Atrous Separable Convolution for Semantic Image Segmentation. Computer Vision – ECCV 2018, Springer International Publishing, Cham, 833-851, (2018).
- [8] Liu Z, Lin Y, Cao Y, Hu H, Wei Y, Zhang Z, et al. Swin Transformer: Hierarchical Vision Transformer using Shifted Windows. 2021 IEEE/CVF International Conference on Computer Vision (ICCV), 9992-10002, (2021).
Ayrıntılar
Birincil Dil
İngilizce
Konular
İnşaat Yapım Mühendisliği, İnşaat Mühendisliği (Diğer), Fotogrametri ve Uzaktan Algılama
Bölüm
Araştırma Makalesi
Erken Görünüm Tarihi
18 Eylül 2026
Yayımlanma Tarihi
-
Gönderilme Tarihi
24 Ağustos 2026
Kabul Tarihi
7 Eylül 2026
Yayımlandığı Sayı
Yıl 2026 Sayı: Advanced Online Publication
