跳到主要导航 跳到搜索 跳到主要内容

CGFT: Cross-Guided Fusion Transformer for Multi-Modal Multi-View Scene Recognition

  • Northwestern Polytechnical University Xian
  • Xi'an University of Science and Technology

科研成果: 期刊稿件文章同行评审

摘要

Multi-modal learning aims to learn comprehensive representations by integrating complementary information from different perspectives. However, multi-modal images acquired from heterogeneous sensor systems frequently exhibit substantial representation gaps, making multi-modal fusion a challenging task. Existing methods often struggle to model semantic correspondences across different modalities, resulting in feature misalignment that hinders effective multi-modal representation learning. To address this issue, we propose a Cross-Guided Fusion Transformer (CGFT) for multi-modal multi-view remote sensing scene recognition. CGFT improves feature alignment through a progressive coarse-to-fine refinement of cross-modal correspondences. Specifically, we first bridge the representation discrepancy by comparing features across modalities with an instance-level contrastive learning loss. Then, we propose a fine-grained alignment module with a sophisticated cross-attention mechanism that explicitly refines semantic correspondences by aligning relevant tokens with cross-modal query-focused guidance. Moreover, we adopt a multi-modal feature fusion scheme with modality-wise weighting to enhance the complementary integration of multi-modal features. This strategy effectively promotes the fine-grained cross-modal alignment while suppressing noisy correspondence problem, thereby leading to enhanced performance in complex real-world scenarios. Extensive experiments on benchmark datasets validate the effectiveness of the proposed method.

学术指纹

探究 'CGFT: Cross-Guided Fusion Transformer for Multi-Modal Multi-View Scene Recognition' 的科研主题。它们共同构成独一无二的学术指纹。

引用此