跳到主要导航 跳到搜索 跳到主要内容

Camouflaged Target Detection Based on Cross Modal Frequency-Domain Feature Enhancement and Text-Driven Fusion Network

  • Northwestern Polytechnical University Xian
  • College of Information and Communication Engineering, Harbin Engineering University

科研成果: 期刊稿件文章同行评审

摘要

Camouflaged target detection (CTD) is a challenging task that involves identifying targets with high visual similarity to their surroundings. Existing methods primarily focus on spatial-domain features, while cross-modal fusion often lacks guidance from high-level semantics, leading to features that are not sufficiently discriminative and are prone to noise interference. To address these issues, we proposes an innovative cross modal frequency-domain feature enhancement and text-driven fusion network (CMFT). First, high-level semantic information is extracted from descriptive text by a frozen CLIP text encoder to provide precise semantic guidance for subsequent fusion. Next, a frequency-domain collaborative enhancement (FDC) module is constructed to enhance high-frequency details in shallow layers and compensate for low-frequency global information in deep layers, thereby improving both the target's detailed texture and global structure in the feature space. Then, a text-driven coupled affine fusion (TCF) module is designed to dynamically generate scaling and translation parameters based on text embeddings, achieving adaptive alignment and fusion of visible and infrared features. Finally, a quality-aware intersection over union loss is introduced to provide sample quality-aware supervision for bounding-box regression, thereby strengthening the learning of reliable predictions and reducing interference from low-quality samples. Experimental results show that CMFT outperforms existing methods on both custom and public benchmark datasets.

指纹

探究 'Camouflaged Target Detection Based on Cross Modal Frequency-Domain Feature Enhancement and Text-Driven Fusion Network' 的科研主题。它们共同构成独一无二的指纹。

引用此