Abstract
Vision Foundation Models (VFMs) pre-trained on large-scale datasets have significantly improved performance in remote sensing semantic segmentation. However, existing methods typically rely on full fine-tuning, which requires updating all model parameters. Instead of updating the full parameter set, Parameter-Efficient Fine-Tuning (PEFT) achieves competitive performance by optimizing only a small subset of parameters. Despite its success, most existing PEFT methods are mainly designed for natural image tasks and fail to account for the unique multi-scale characteristics of remote sensing images. To address these challenges, we propose Multi-scale Cognitive Feature Refinement Tuning, named MsRE, a novel PEFT method tailored for remote sensing semantic segmentation. Specifically, MsRE captures multi-scale contextual information by applying cognitive operations with different cognitive fields to intermediate features of backbone. It then introduces a set of learnable tokens to establish interactions with features at different scales, enabling precise feature refinement and progressive feature propagation across network layers. This mechanism enhances the model’s ability to understand complex remote sensing scenes and improves downstream segmentation performance. With significantly fewer trainable parameters, MsRE provides an efficient yet effective solution for adapting VFMs to remote sensing segmentation tasks. Extensive experiments demonstrate that MsRE achieves competitive segmentation performance with substantially fewer trainable backbone parameters, providing a favorable balance between accuracy and parameter efficiency. The project is available at http://woldier.top/MsRE.
Performance Highlights
Semantic segmentation result (mIoU%) comparisons across different methods on the Potsdam, Vaihingen, and LoveDA datasets.
| Method | Pre-TrainMethod | Pre-TrainDataset | Backboneand Scale | Fine-TuneParams↓ | Potsdam(mIoU%) | Vaihigen(mIoU%) | LoveDA(mIoU%) | Average(mIoU%) |
|---|---|---|---|---|---|---|---|---|
| I. Specific Segmentation Methods | ||||||||
| 1. General Segmentation Methods | ||||||||
| DeepLabV3+ | Supervised | ImageNet1K | Resnet-50 | 25.2 M | 73.43 | 71.87 | 45.33 | 63.54 |
| UperNet | Supervised | ImageNet1K | Resnet-50 | 25.2 M | 73.43 | 71.86 | 46.28 | 63.85 |
| Segformer | Supervised | ImageNet1K | MiT-B5 | 89.9 M | 79.01 | 74.38 | 49.69 | 67.69 |
| Segmenter | Supervised | ImageNet1K | ViT-Large | 304.3 M | 78.24 | 73.12 | 53.34 | 68.23 |
| 2. Remote Sensing Segmentation Methods | ||||||||
| UNetFormer | Supervised | ImageNet1K | ResNet-18 | 11.6 M | 74.16 | 72.38 | 46.83 | 64.45 |
| FSegNet | Supervised | ImageNet1K | FasterViT-3 | 159.5 M | 78.17 | 74.51 | 48.93 | 67.20 |
| MSEONet | Supervised | ImageNet1K | Resnet-101 | 44.5 M | 74.51 | 70.08 | 45.17 | 63.25 |
| EMGSNet | Supervised | ImageNet1K | PVT-Base | 42.1 M | 79.64 | 75.84 | 52.95 | 69.47 |
| II. Vision Foundation Models | ||||||||
| Full Fine-Tuning | DINOv3 | LVD-1689M | ViT-Large | 303.1 M | 81.04 | 75.12 | 54.61 | 70.25 |
| Frozen | DINOv3 | LVD-1689M | ViT-Large | 0 M | 79.21 | 74.34 | 52.55 | 68.70 |
| i. General PEFT Methods | ||||||||
| LoRA | DINOv3 | LVD-1689M | ViT-Large | 9.43 M | 79.67 | 73.25 | 53.19 | 68.70 |
| Adapter | DINOv3 | LVD-1689M | ViT-Large | 6.24 M | 80.35 | 73.91 | 54.18 | 69.48 |
| AdaptFormer | DINOv3 | LVD-1689M | ViT-Large | 3.12 M | 79.34 | 76.12 | 15.18 | 56.88 |
| Mona | DINOv3 | LVD-1689M | ViT-Large | 5.08 M | 80.39 | 74.96 | 53.87 | 69.74 |
| ii. Remote Sensing PEFT Methods | ||||||||
| AiRs | DINOv3 | LVD-1689M | ViT-Large | 4.24 M | 79.94 | 74.78 | 54.11 | 69.61 |
| Tea | DINOv3 | LVD-1689M | ViT-Large | 11.32 M | 80.43 | 75.01 | 54.36 | 69.93 |
| iii. Ours Approach | ||||||||
| MsRE (ours) | DINOv3 | LVD-1689M | ViT-Large | 4.64 M | 80.98 | 76.72 | 54.90 | 70.86 |
Notes:
- The results in each column is highlighted in bold.
- All Vision Foundation models are evaluated using the Segmenter head for fair comparison.
- "Fine-Tune Params" refer to the number of trainable parameters in the backbone on the downstream tasks.
- "Supervised" stands for supervised learning.
Experiment Results && Visualization
Class-wise IoU comparison on Potsdam, Vaihingen and LoveDA.
Visualization of results on Potsdam. Note that all methods are based on DINOv3
Visualization of results on LoveDA. Note that all methods are based on DINOv3.
PCA visualization of feature representations from different layers (Layer 6, 12, 18, and 24) of MsRE on the Potsdam dataset.
BibTeX
@ARTICLE{11599658,
author={Wang, Bin and Lv, Shun and Li, Zhi and Deng, Fei and Liu, Yiguang},
journal={IEEE Transactions on Geoscience and Remote Sensing},
title={MsRE: Towards Efficient Remote Sensing Segmentation via Vision Foundation Models},
year={2026},
volume={},
number={},
pages={1-1},
keywords={Modeling;Remote sensing;Semantic segmentation;Training;Tuning;Decoding;LoRa;Visualization;Vegetation;Head;Vision Foundation Models;Semantic Segmentation;Parameter-Efficient Fine-Tuning;Remote Sensing},
doi={10.1109/TGRS.2026.3711219}
}