TY - JOUR
T1 - TPTAF
T2 - Task-Prior Tripartite Attention for Infrared and Visible Image Fusion
AU - Du, Xuyang
AU - Yao, Xiwen
AU - Zang, Ankang
AU - Cheng, Gong
N1 - Publisher Copyright:
© 1992-2012 IEEE.
PY - 2026
Y1 - 2026
N2 - Infrared and visible image fusion aims to integrate complementary information to produce informative fused images for visual perception and downstream tasks. However, mainstream methods often focus on fusion reconstruction, while losses from different tasks may introduce competing optimization demands, making it difficult to balance visual quality and detection performance. To address this issue, we propose task-prior tripartite attention for infrared and visible image fusion (TPTAF), which introduces detection semantics as task-prior guidance for representation-level cross-modal interaction. TPTAF employs a decoupled encoder to organize infrared and visible features into structural and discriminative spaces, enabling cross-modal layout consistency and modality-specific detail preservation to be modeled with different roles. Meanwhile, detection semantics extracted from hybrid infrared-visible features are transformed into task priors and integrated with the decoupled representations through tripartite attention. In this interaction, structural cues stabilize the fusion layout, while task-prior guidance regulates the selection of discriminative details toward target-related evidence. For joint optimization, uncertainty-weighted learning further balances fusion and detection losses, reducing the dependence on manually assigned loss weights. Experiments on M3FD, Road-Scene, AVMS, and MSRS demonstrate that TPTAF maintains stable fusion quality across different data distributions while improving the utility of fused images for downstream object detection and semantic segmentation evaluation.
AB - Infrared and visible image fusion aims to integrate complementary information to produce informative fused images for visual perception and downstream tasks. However, mainstream methods often focus on fusion reconstruction, while losses from different tasks may introduce competing optimization demands, making it difficult to balance visual quality and detection performance. To address this issue, we propose task-prior tripartite attention for infrared and visible image fusion (TPTAF), which introduces detection semantics as task-prior guidance for representation-level cross-modal interaction. TPTAF employs a decoupled encoder to organize infrared and visible features into structural and discriminative spaces, enabling cross-modal layout consistency and modality-specific detail preservation to be modeled with different roles. Meanwhile, detection semantics extracted from hybrid infrared-visible features are transformed into task priors and integrated with the decoupled representations through tripartite attention. In this interaction, structural cues stabilize the fusion layout, while task-prior guidance regulates the selection of discriminative details toward target-related evidence. For joint optimization, uncertainty-weighted learning further balances fusion and detection losses, reducing the dependence on manually assigned loss weights. Experiments on M3FD, Road-Scene, AVMS, and MSRS demonstrate that TPTAF maintains stable fusion quality across different data distributions while improving the utility of fused images for downstream object detection and semantic segmentation evaluation.
KW - Attention module
KW - feature encoder
KW - image fusion
KW - object detection
KW - uncertainty-weighted learning
UR - https://www.scopus.com/pages/publications/105044724899
U2 - 10.1109/TIP.2026.3709518
DO - 10.1109/TIP.2026.3709518
M3 - 文章
C2 - 42424210
AN - SCOPUS:105044724899
SN - 1057-7149
JO - IEEE Transactions on Image Processing
JF - IEEE Transactions on Image Processing
ER -