MsRE: Towards Efficient Remote Sensing Segmentation via Vision Foundation Models

Bin Wang1, Shun Lv1, Zhi Li1, Fei Deng2, Yiguang Liu1
1 Sichuan University 2 Chengdu University of Technology
TGRS 2026
Overview

Overview of MsRE.

Abstract

Vision Foundation Models (VFMs) pre-trained on large-scale datasets have significantly improved performance in remote sensing semantic segmentation. However, existing methods typically rely on full fine-tuning, which requires updating all model parameters. Instead of updating the full parameter set, Parameter-Efficient Fine-Tuning (PEFT) achieves competitive performance by optimizing only a small subset of parameters. Despite its success, most existing PEFT methods are mainly designed for natural image tasks and fail to account for the unique multi-scale characteristics of remote sensing images. To address these challenges, we propose Multi-scale Cognitive Feature Refinement Tuning, named MsRE, a novel PEFT method tailored for remote sensing semantic segmentation. Specifically, MsRE captures multi-scale contextual information by applying cognitive operations with different cognitive fields to intermediate features of backbone. It then introduces a set of learnable tokens to establish interactions with features at different scales, enabling precise feature refinement and progressive feature propagation across network layers. This mechanism enhances the model’s ability to understand complex remote sensing scenes and improves downstream segmentation performance. With significantly fewer trainable parameters, MsRE provides an efficient yet effective solution for adapting VFMs to remote sensing segmentation tasks. Extensive experiments demonstrate that MsRE achieves competitive segmentation performance with substantially fewer trainable backbone parameters, providing a favorable balance between accuracy and parameter efficiency. The project is available at http://woldier.top/MsRE.

Performance Highlights

Semantic segmentation result (mIoU%) comparisons across different methods on the Potsdam, Vaihingen, and LoveDA datasets.

Method Pre-Train
Method
Pre-Train
Dataset
Backbone
and Scale
Fine-Tune
Params
Potsdam
(mIoU%)
Vaihigen
(mIoU%)
LoveDA
(mIoU%)
Average
(mIoU%)
I. Specific Segmentation Methods
1. General Segmentation Methods
DeepLabV3+ Supervised ImageNet1K Resnet-50 25.2 M 73.43 71.87 45.33 63.54
UperNet Supervised ImageNet1K Resnet-50 25.2 M 73.43 71.86 46.28 63.85
Segformer Supervised ImageNet1K MiT-B5 89.9 M 79.01 74.38 49.69 67.69
Segmenter Supervised ImageNet1K ViT-Large 304.3 M 78.24 73.12 53.34 68.23
2. Remote Sensing Segmentation Methods
UNetFormer Supervised ImageNet1K ResNet-18 11.6 M 74.16 72.38 46.83 64.45
FSegNet Supervised ImageNet1K FasterViT-3 159.5 M 78.17 74.51 48.93 67.20
MSEONet Supervised ImageNet1K Resnet-101 44.5 M 74.51 70.08 45.17 63.25
EMGSNet Supervised ImageNet1K PVT-Base 42.1 M 79.64 75.84 52.95 69.47
II. Vision Foundation Models
Full Fine-Tuning DINOv3 LVD-1689M ViT-Large 303.1 M 81.04 75.12 54.61 70.25
Frozen DINOv3 LVD-1689M ViT-Large 0 M 79.21 74.34 52.55 68.70
i. General PEFT Methods
LoRA DINOv3 LVD-1689M ViT-Large 9.43 M 79.67 73.25 53.19 68.70
Adapter DINOv3 LVD-1689M ViT-Large 6.24 M 80.35 73.91 54.18 69.48
AdaptFormer DINOv3 LVD-1689M ViT-Large 3.12 M 79.34 76.12 15.18 56.88
Mona DINOv3 LVD-1689M ViT-Large 5.08 M 80.39 74.96 53.87 69.74
ii. Remote Sensing PEFT Methods
AiRs DINOv3 LVD-1689M ViT-Large 4.24 M 79.94 74.78 54.11 69.61
Tea DINOv3 LVD-1689M ViT-Large 11.32 M 80.43 75.01 54.36 69.93
iii. Ours Approach
MsRE (ours) DINOv3 LVD-1689M ViT-Large 4.64 M 80.98 76.72 54.90 70.86

Notes:

  • The results in each column is highlighted in bold.
  • All Vision Foundation models are evaluated using the Segmenter head for fair comparison.
  • "Fine-Tune Params" refer to the number of trainable parameters in the backbone on the downstream tasks.
  • "Supervised" stands for supervised learning.

Experiment Results && Visualization

BibTeX


        @ARTICLE{11599658,
          author={Wang, Bin and Lv, Shun and Li, Zhi and Deng, Fei and Liu, Yiguang},
          journal={IEEE Transactions on Geoscience and Remote Sensing},
          title={MsRE: Towards Efficient Remote Sensing Segmentation via Vision Foundation Models},
          year={2026},
          volume={},
          number={},
          pages={1-1},
          keywords={Modeling;Remote sensing;Semantic segmentation;Training;Tuning;Decoding;LoRa;Visualization;Vegetation;Head;Vision Foundation Models;Semantic Segmentation;Parameter-Efficient Fine-Tuning;Remote Sensing},
          doi={10.1109/TGRS.2026.3711219}
        }