TY - JOUR
T1 - A Two-Stage Unified Refinement Framework with RS-Specific Priors for Remote Sensing Open-Vocabulary Segmentation
AU - Li, Jiayuan
AU - Wang, Zhen
AU - Sun, Xiao
AU - Li, Yue Chao
AU - Huang, Yu An
AU - Zhang, Zhe
AU - You, Zhu Hong
N1 - Publisher Copyright:
© 1980-2012 IEEE.
PY - 2026
Y1 - 2026
N2 - Vision-language pre-trained models, particularly CLIP, have demonstrated remarkable zero-shot transfer capabilities across various image-level tasks, catalyzing the advancement of open-vocabulary semantic segmentation (OVSS) in remote sensing (RS). However, the direct deployment of CLIP to the RS domain is inherently constrained by the profound domain shift between terrestrial and overhead perspectives, as well as the intricate geometric heterogeneities regarding scale and orientation. To circumvent these limitations, we propose CDSeg, a robust framework tailored for RSOVSS. Central to this architecture is the Dual-Domain Feature Compensation Module (DDFCM), which integrates DINOv3 weights, pretrained on large-scale RS benchmarks, to augment CLIP with domain-specific semantic priors, effectively bridging the natural-to-satellite knowledge gap. Furthermore, we introduce a MambaVision-driven Cross Feature Fine-grained Interaction Module (CFFIM) to facilitate a unified refinement of spatial and category attributes, leveraging long-range dependency modeling to enhance the model’s discriminative power in unseen environments. To robustly manage the complexities of diverse orientations and scales, CDSeg incorporates a direction-aware rotation strategy and a Wavelet-Cross-Attention Enhanced Module (WCAEM) for high-fidelity multiscale feature decoding. Empirical evaluations on four public benchmarks demonstrate that CDSeg achieves state-of-the-art performance, while extensive ablation studies substantiate the synergistic contribution and indispensability of each component.
AB - Vision-language pre-trained models, particularly CLIP, have demonstrated remarkable zero-shot transfer capabilities across various image-level tasks, catalyzing the advancement of open-vocabulary semantic segmentation (OVSS) in remote sensing (RS). However, the direct deployment of CLIP to the RS domain is inherently constrained by the profound domain shift between terrestrial and overhead perspectives, as well as the intricate geometric heterogeneities regarding scale and orientation. To circumvent these limitations, we propose CDSeg, a robust framework tailored for RSOVSS. Central to this architecture is the Dual-Domain Feature Compensation Module (DDFCM), which integrates DINOv3 weights, pretrained on large-scale RS benchmarks, to augment CLIP with domain-specific semantic priors, effectively bridging the natural-to-satellite knowledge gap. Furthermore, we introduce a MambaVision-driven Cross Feature Fine-grained Interaction Module (CFFIM) to facilitate a unified refinement of spatial and category attributes, leveraging long-range dependency modeling to enhance the model’s discriminative power in unseen environments. To robustly manage the complexities of diverse orientations and scales, CDSeg incorporates a direction-aware rotation strategy and a Wavelet-Cross-Attention Enhanced Module (WCAEM) for high-fidelity multiscale feature decoding. Empirical evaluations on four public benchmarks demonstrate that CDSeg achieves state-of-the-art performance, while extensive ablation studies substantiate the synergistic contribution and indispensability of each component.
KW - CDSeg
KW - CFFIM
KW - DDFCM
KW - Remote Sensing Open-Vocabulary Segmentation
KW - WCAEM
UR - https://www.scopus.com/pages/publications/105044359038
U2 - 10.1109/TGRS.2026.3710833
DO - 10.1109/TGRS.2026.3710833
M3 - 文章
AN - SCOPUS:105044359038
SN - 0196-2892
JO - IEEE Transactions on Geoscience and Remote Sensing
JF - IEEE Transactions on Geoscience and Remote Sensing
ER -