TY - JOUR
T1 - Not All Encoder Layers Are Essential
T2 - Shallow-Feature Fusion for Infrared Small Target Detection
AU - Li, Qiang
AU - Zheng, Jiangbin
AU - Zou, Wenbin
AU - Zhao, Yong
AU - Wang, Bingshu
AU - Zhang, Han
N1 - Publisher Copyright:
© 1980-2012 IEEE.
PY - 2026
Y1 - 2026
N2 - Deep learning has revolutionized infrared small target detection. However, prevailing encoder-decoder architectures, which maintain symmetrical depths, incur unnecessary computational cost. We challenge this design with a key insight: for this task, deep encoder layers are largely dispensable in the decoder feature fusion process, as shallow features alone can achieve an optimal accuracy-efficiency balance. Driven by this insight, we propose TADNet, a Target-Aware Dual-branch Network, to translate this insight into an efficient and accurate model. Its core components incorporate: (1) an effective dual-branch encoder that collaboratively extracts high-resolution spatial details and rich semantic contexts, and (2) an efficient dynamic fusion decoder that strategically utilizes only the first two shallowest features— pruning the traditional fusion backbone. Beyond the model, we introduce the NPU-SIRST benchmark to address multi-scale and multi-scenario limitations. Rigorous experiments on four datasets show that our method achieves state-of-the-art results with significantly enhanced efficiency, demonstrating robust transferability under infrared cross-dataset testing.
AB - Deep learning has revolutionized infrared small target detection. However, prevailing encoder-decoder architectures, which maintain symmetrical depths, incur unnecessary computational cost. We challenge this design with a key insight: for this task, deep encoder layers are largely dispensable in the decoder feature fusion process, as shallow features alone can achieve an optimal accuracy-efficiency balance. Driven by this insight, we propose TADNet, a Target-Aware Dual-branch Network, to translate this insight into an efficient and accurate model. Its core components incorporate: (1) an effective dual-branch encoder that collaboratively extracts high-resolution spatial details and rich semantic contexts, and (2) an efficient dynamic fusion decoder that strategically utilizes only the first two shallowest features— pruning the traditional fusion backbone. Beyond the model, we introduce the NPU-SIRST benchmark to address multi-scale and multi-scenario limitations. Rigorous experiments on four datasets show that our method achieves state-of-the-art results with significantly enhanced efficiency, demonstrating robust transferability under infrared cross-dataset testing.
KW - deep learning
KW - encoder-decoder
KW - Infrared small target detection
KW - shallow feature
UR - https://www.scopus.com/pages/publications/105041985574
U2 - 10.1109/TGRS.2026.3702817
DO - 10.1109/TGRS.2026.3702817
M3 - 文章
AN - SCOPUS:105041985574
SN - 0196-2892
JO - IEEE Transactions on Geoscience and Remote Sensing
JF - IEEE Transactions on Geoscience and Remote Sensing
ER -