Abstract
Deep learning has revolutionized infrared small target detection. However, prevailing encoder-decoder architectures, which maintain symmetrical depths, incur unnecessary computational cost. We challenge this design with a key insight: for this task, deep encoder layers are largely dispensable in the decoder feature fusion process, as shallow features alone can achieve an optimal accuracy-efficiency balance. Driven by this insight, we propose TADNet, a Target-Aware Dual-branch Network, to translate this insight into an efficient and accurate model. Its core components incorporate: (1) an effective dual-branch encoder that collaboratively extracts high-resolution spatial details and rich semantic contexts, and (2) an efficient dynamic fusion decoder that strategically utilizes only the first two shallowest features— pruning the traditional fusion backbone. Beyond the model, we introduce the NPU-SIRST benchmark to address multi-scale and multi-scenario limitations. Rigorous experiments on four datasets show that our method achieves state-of-the-art results with significantly enhanced efficiency, demonstrating robust transferability under infrared cross-dataset testing.
| Original language | English |
|---|---|
| Journal | IEEE Transactions on Geoscience and Remote Sensing |
| DOIs | |
| State | Accepted/In press - 2026 |
Keywords
- deep learning
- encoder-decoder
- Infrared small target detection
- shallow feature
Fingerprint
Dive into the research topics of 'Not All Encoder Layers Are Essential: Shallow-Feature Fusion for Infrared Small Target Detection'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver