Abstract
Leveraging the complementary characteristics of infrared and visible images enables the construction of more comprehensive scene representations and improves performance across diverse downstream vision tasks. However, most existing approaches treat various infrared-visible (IR-VIS) vision tasks (e.g., image fusion, object detection, semantic segmentation, and crowd counting) as independent problems, typically employing separate models for each. This paradigm not only results in increased computational and model redundancy but also hinders the exploitation of potential synergies among tasks. To address this limitation, we propose MTFusion, a unified multi-task vision framework designed for joint infrared-visible image fusion (IVIF) and general vision applications. The core idea of MTFusion lies in achieving synergistic collaboration among IR-VIS vision tasks through a “commonality-and-specificity” modeling strategy, that is, by jointly exploring shared semantic foundations while preserving the distinct characteristics of each task. Specifically, MTFusion integrates low-level fusion and multiple high-level tasks within a single architecture, enabling shared access to multiscale, cross-modal fused representations. On the one hand, we leverage the powerful semantic representation capability of a foundation model to enrich the multimodal backbone with common semantic cues, ensuring consistent understanding across tasks. On the other hand, we introduce task-specific textual prompts to guide feature projection into individualized subspaces, imposing semantic priors that enhance task adaptability and discrimination. Extensive experiments on multiple benchmark datasets demonstrate that MTFusion effectively balances shared representation learning and task-specific specialization, achieving state-of-the-art performance across various IR-VIS vision tasks.
| Original language | English |
|---|---|
| Article number | 133865 |
| Journal | Expert Systems with Applications |
| Volume | 333 |
| DOIs | |
| State | Published - 1 Jan 2027 |
Keywords
- Commonality-and-specificity modeling
- Infrared-visible image fusion
- Multi-task adaptation
- Multi-task learning
Fingerprint
Dive into the research topics of 'MTFusion: A unified multi-task framework for joint infrared-visible image fusion and general vision tasks'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver