TY - JOUR
T1 - Camouflaged Target Detection Based on Cross Modal Frequency-Domain Feature Enhancement and Text-Driven Fusion Network
AU - Lai, Jie
AU - Geng, Jie
AU - Peng, Ruihui
AU - Jiang, Wen
N1 - Publisher Copyright:
© 1965-2011 IEEE.
PY - 2026
Y1 - 2026
N2 - Camouflaged target detection (CTD) is a challenging task that involves identifying targets with high visual similarity to their surroundings. Existing methods primarily focus on spatial-domain features, while cross-modal fusion often lacks guidance from high-level semantics, leading to features that are not sufficiently discriminative and are prone to noise interference. To address these issues, we proposes an innovative cross modal frequency-domain feature enhancement and text-driven fusion network (CMFT). First, high-level semantic information is extracted from descriptive text by a frozen CLIP text encoder to provide precise semantic guidance for subsequent fusion. Next, a frequency-domain collaborative enhancement (FDC) module is constructed to enhance high-frequency details in shallow layers and compensate for low-frequency global information in deep layers, thereby improving both the target's detailed texture and global structure in the feature space. Then, a text-driven coupled affine fusion (TCF) module is designed to dynamically generate scaling and translation parameters based on text embeddings, achieving adaptive alignment and fusion of visible and infrared features. Finally, a quality-aware intersection over union loss is introduced to provide sample quality-aware supervision for bounding-box regression, thereby strengthening the learning of reliable predictions and reducing interference from low-quality samples. Experimental results show that CMFT outperforms existing methods on both custom and public benchmark datasets.
AB - Camouflaged target detection (CTD) is a challenging task that involves identifying targets with high visual similarity to their surroundings. Existing methods primarily focus on spatial-domain features, while cross-modal fusion often lacks guidance from high-level semantics, leading to features that are not sufficiently discriminative and are prone to noise interference. To address these issues, we proposes an innovative cross modal frequency-domain feature enhancement and text-driven fusion network (CMFT). First, high-level semantic information is extracted from descriptive text by a frozen CLIP text encoder to provide precise semantic guidance for subsequent fusion. Next, a frequency-domain collaborative enhancement (FDC) module is constructed to enhance high-frequency details in shallow layers and compensate for low-frequency global information in deep layers, thereby improving both the target's detailed texture and global structure in the feature space. Then, a text-driven coupled affine fusion (TCF) module is designed to dynamically generate scaling and translation parameters based on text embeddings, achieving adaptive alignment and fusion of visible and infrared features. Finally, a quality-aware intersection over union loss is introduced to provide sample quality-aware supervision for bounding-box regression, thereby strengthening the learning of reliable predictions and reducing interference from low-quality samples. Experimental results show that CMFT outperforms existing methods on both custom and public benchmark datasets.
KW - Camouflaged target detection
KW - frequency domain enhancement
KW - sample quality perception
KW - text semantic guidance
UR - https://www.scopus.com/pages/publications/105041082337
U2 - 10.1109/TAES.2026.3699317
DO - 10.1109/TAES.2026.3699317
M3 - 文章
AN - SCOPUS:105041082337
SN - 0018-9251
JO - IEEE Transactions on Aerospace and Electronic Systems
JF - IEEE Transactions on Aerospace and Electronic Systems
ER -