摘要
Multi-modality image fusion aims to integrate the merits of images from different sources and render high-quality fused images. Existing CNN-based image fusion backbones are limited by biases caused by local receptive fields and static parameters during inference. Transformer-based models have global receptive fields, but to avoid the trouble caused by quadratic computational complexity, they usually use self-attention mechanisms with fixed range windows or along the channel dimension, which limits the utilization of the global receptive field. To address this problem, we propose a hybrid Transformer-Mamba structure, called TMamba. Mamba has excellent global feature extraction capabilities but lacks efficient interactions between its channels. We use the linear Transformer to capture features through token interactions between channels and use the Mamba to capture features through token interactions at different spatial locations. It allows the model to extract global features from different dimensions and optimize each other while maintaining linear complexity. We further introduce three-level interactions among branches, modalities, and TMamba structures. Branch interaction enables channel and spatial global features to be transferred between Transformer and Mamba branches. Modality interaction integrates features of different modalities. TMamba interaction integrates and strengthens features from different branches. Experiments show that our TMamba achieves excellent results in multiple fusion tasks, including infrared-visible image fusion and medical image fusion. Code with checkpoints will be provided after peer review.
| 源语言 | 英语 |
|---|---|
| 期刊 | IEEE Transactions on Circuits and Systems for Video Technology |
| DOI | |
| 出版状态 | 已接受/待刊 - 2026 |
指纹
探究 'TMamba: Global Channel-Location Token Interaction for Multi-Modality Image Fusion' 的科研主题。它们共同构成独一无二的指纹。引用此
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver