Abstract
Recent progress in prompt-driven vision transformers has markedly enhanced remote sensing image (RSI) semantic segmentation. However, existing methods often fail to incorporate fine-grained geometric structure and to effectively address the spatial heterogeneity present in complex aerial scenes. To overcome these limitations, we propose GeoMoE, a geometry-driven adaptive mixture-of-experts framework tailored for precise and efficient semantic segmentation. Building upon a frozen vision transformer backbone, GeoMoE introduces three novel components, 1) a token-field hybrid adapter (TFH-Adapter) that enables geometry-aware feature adaptation without modifying backbone parameters; 2) a geometry-contextual prompt generator (Geo-Prompt) that integrates multi-scale shape prototypes and contextual cues into expressive prompt embeddings; and 3) a geometry mixture-of-complexity-experts (Geo-MoCE) decoder which dynamically routes spatial regions to specialized experts based on local geometric complexity. This unified architecture allows explicit modeling of geometric information and flexible allocation of decoding capacity, resulting in more accurate segmentation of structurally complex and heterogeneous regions. Extensive experiments on several benchmark remote sensing datasets demonstrate that GeoMoE achieves state-of-the-art performance in segmentation accuracy, model efficiency, and boundary delineation.
| Original language | English |
|---|---|
| Article number | 133600 |
| Journal | Expert Systems with Applications |
| Volume | 332 |
| DOIs | |
| State | Published - 1 Jan 2027 |
Keywords
- Mixture-of-Experts
- Prompt generation
- Remote sensing
- Semantic segmentation
- Vision transformer
Fingerprint
Dive into the research topics of 'GeoMoE: Geometry-Driven prompts with adaptive Mixture-of-Experts for remote sensing image semantic segmentation'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver