跳到主要导航 跳到搜索 跳到主要内容

MTFusion: A unified multi-task framework for joint infrared-visible image fusion and general vision tasks

  • Northwestern Polytechnical University Xian
  • Xi'an Institute of Posts and Telecommunications

科研成果: 期刊稿件文章同行评审

摘要

Leveraging the complementary characteristics of infrared and visible images enables the construction of more comprehensive scene representations and improves performance across diverse downstream vision tasks. However, most existing approaches treat various infrared-visible (IR-VIS) vision tasks (e.g., image fusion, object detection, semantic segmentation, and crowd counting) as independent problems, typically employing separate models for each. This paradigm not only results in increased computational and model redundancy but also hinders the exploitation of potential synergies among tasks. To address this limitation, we propose MTFusion, a unified multi-task vision framework designed for joint infrared-visible image fusion (IVIF) and general vision applications. The core idea of MTFusion lies in achieving synergistic collaboration among IR-VIS vision tasks through a “commonality-and-specificity” modeling strategy, that is, by jointly exploring shared semantic foundations while preserving the distinct characteristics of each task. Specifically, MTFusion integrates low-level fusion and multiple high-level tasks within a single architecture, enabling shared access to multiscale, cross-modal fused representations. On the one hand, we leverage the powerful semantic representation capability of a foundation model to enrich the multimodal backbone with common semantic cues, ensuring consistent understanding across tasks. On the other hand, we introduce task-specific textual prompts to guide feature projection into individualized subspaces, imposing semantic priors that enhance task adaptability and discrimination. Extensive experiments on multiple benchmark datasets demonstrate that MTFusion effectively balances shared representation learning and task-specific specialization, achieving state-of-the-art performance across various IR-VIS vision tasks.

源语言英语
期刊论文编号133865
期刊Expert Systems with Applications
333
DOI
出版状态已出版 - 1 1月 2027

学术指纹

探究 'MTFusion: A unified multi-task framework for joint infrared-visible image fusion and general vision tasks' 的科研主题。它们共同构成独一无二的学术指纹。

引用此