Abstract
Multi-modal learning aims to learn comprehensive representations by integrating complementary information from different perspectives. However, multi-modal images acquired from heterogeneous sensor systems frequently exhibit substantial representation gaps, making multi-modal fusion a challenging task. Existing methods often struggle to model semantic correspondences across different modalities, resulting in feature misalignment that hinders effective multi-modal representation learning. To address this issue, we propose a Cross-Guided Fusion Transformer (CGFT) for multi-modal multi-view remote sensing scene recognition. CGFT improves feature alignment through a progressive coarse-to-fine refinement of cross-modal correspondences. Specifically, we first bridge the representation discrepancy by comparing features across modalities with an instance-level contrastive learning loss. Then, we propose a fine-grained alignment module with a sophisticated cross-attention mechanism that explicitly refines semantic correspondences by aligning relevant tokens with cross-modal query-focused guidance. Moreover, we adopt a multi-modal feature fusion scheme with modality-wise weighting to enhance the complementary integration of multi-modal features. This strategy effectively promotes the fine-grained cross-modal alignment while suppressing noisy correspondence problem, thereby leading to enhanced performance in complex real-world scenarios. Extensive experiments on benchmark datasets validate the effectiveness of the proposed method.
| Original language | English |
|---|---|
| Journal | IEEE Transactions on Geoscience and Remote Sensing |
| DOIs | |
| State | Accepted/In press - 2026 |
Keywords
- cross-guided
- information fusion
- multi-modal
- multi-view
- Scene recognition
- vision transformer
Fingerprint
Dive into the research topics of 'CGFT: Cross-Guided Fusion Transformer for Multi-Modal Multi-View Scene Recognition'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver