Abstract
Camouflaged target detection (CTD) is a challenging task that involves identifying targets with high visual similarity to their surroundings. Existing methods primarily focus on spatial-domain features, while cross-modal fusion often lacks guidance from high-level semantics, leading to features that are not sufficiently discriminative and are prone to noise interference. To address these issues, we proposes an innovative cross modal frequency-domain feature enhancement and text-driven fusion network (CMFT). First, high-level semantic information is extracted from descriptive text by a frozen CLIP text encoder to provide precise semantic guidance for subsequent fusion. Next, a frequency-domain collaborative enhancement (FDC) module is constructed to enhance high-frequency details in shallow layers and compensate for low-frequency global information in deep layers, thereby improving both the target's detailed texture and global structure in the feature space. Then, a text-driven coupled affine fusion (TCF) module is designed to dynamically generate scaling and translation parameters based on text embeddings, achieving adaptive alignment and fusion of visible and infrared features. Finally, a quality-aware intersection over union loss is introduced to provide sample quality-aware supervision for bounding-box regression, thereby strengthening the learning of reliable predictions and reducing interference from low-quality samples. Experimental results show that CMFT outperforms existing methods on both custom and public benchmark datasets.
| Original language | English |
|---|---|
| Journal | IEEE Transactions on Aerospace and Electronic Systems |
| DOIs | |
| State | Accepted/In press - 2026 |
Keywords
- Camouflaged target detection
- frequency domain enhancement
- sample quality perception
- text semantic guidance
Fingerprint
Dive into the research topics of 'Camouflaged Target Detection Based on Cross Modal Frequency-Domain Feature Enhancement and Text-Driven Fusion Network'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver