Skip to main navigation Skip to search Skip to main content

CGFT: Cross-Guided Fusion Transformer for Multi-Modal Multi-View Scene Recognition

  • Northwestern Polytechnical University Xian
  • Xi'an University of Science and Technology

Research output: Contribution to journalArticlepeer-review

Abstract

Multi-modal learning aims to learn comprehensive representations by integrating complementary information from different perspectives. However, multi-modal images acquired from heterogeneous sensor systems frequently exhibit substantial representation gaps, making multi-modal fusion a challenging task. Existing methods often struggle to model semantic correspondences across different modalities, resulting in feature misalignment that hinders effective multi-modal representation learning. To address this issue, we propose a Cross-Guided Fusion Transformer (CGFT) for multi-modal multi-view remote sensing scene recognition. CGFT improves feature alignment through a progressive coarse-to-fine refinement of cross-modal correspondences. Specifically, we first bridge the representation discrepancy by comparing features across modalities with an instance-level contrastive learning loss. Then, we propose a fine-grained alignment module with a sophisticated cross-attention mechanism that explicitly refines semantic correspondences by aligning relevant tokens with cross-modal query-focused guidance. Moreover, we adopt a multi-modal feature fusion scheme with modality-wise weighting to enhance the complementary integration of multi-modal features. This strategy effectively promotes the fine-grained cross-modal alignment while suppressing noisy correspondence problem, thereby leading to enhanced performance in complex real-world scenarios. Extensive experiments on benchmark datasets validate the effectiveness of the proposed method.

Original languageEnglish
JournalIEEE Transactions on Geoscience and Remote Sensing
DOIs
StateAccepted/In press - 2026

Keywords

  • cross-guided
  • information fusion
  • multi-modal
  • multi-view
  • Scene recognition
  • vision transformer

Fingerprint

Dive into the research topics of 'CGFT: Cross-Guided Fusion Transformer for Multi-Modal Multi-View Scene Recognition'. Together they form a unique fingerprint.

Cite this