摘要
Multi-modal learning aims to learn comprehensive representations by integrating complementary information from different perspectives. However, multi-modal images acquired from heterogeneous sensor systems frequently exhibit substantial representation gaps, making multi-modal fusion a challenging task. Existing methods often struggle to model semantic correspondences across different modalities, resulting in feature misalignment that hinders effective multi-modal representation learning. To address this issue, we propose a Cross-Guided Fusion Transformer (CGFT) for multi-modal multi-view remote sensing scene recognition. CGFT improves feature alignment through a progressive coarse-to-fine refinement of cross-modal correspondences. Specifically, we first bridge the representation discrepancy by comparing features across modalities with an instance-level contrastive learning loss. Then, we propose a fine-grained alignment module with a sophisticated cross-attention mechanism that explicitly refines semantic correspondences by aligning relevant tokens with cross-modal query-focused guidance. Moreover, we adopt a multi-modal feature fusion scheme with modality-wise weighting to enhance the complementary integration of multi-modal features. This strategy effectively promotes the fine-grained cross-modal alignment while suppressing noisy correspondence problem, thereby leading to enhanced performance in complex real-world scenarios. Extensive experiments on benchmark datasets validate the effectiveness of the proposed method.
| 源语言 | 英语 |
|---|---|
| 期刊 | IEEE Transactions on Geoscience and Remote Sensing |
| DOI | |
| 出版状态 | 已接受/待刊 - 2026 |
学术指纹
探究 'CGFT: Cross-Guided Fusion Transformer for Multi-Modal Multi-View Scene Recognition' 的科研主题。它们共同构成独一无二的学术指纹。引用此
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver