摘要
The existing Transformer-based redgreenblue-thermal (RGBT) tracker mainly focuses on the enhancement of features extracted by convolutional neural network (CNN). The potential of the Transformer in representation learning remains underexplored. In this letter, we propose a Convolution-Transformer network with joint multimodal feature learning (JMFL), in which both representation learning and feature fusion leverage Transformer. Specifically, we use the multibranch Convolution-Transformer feature extraction network to process the extraction task of local modality-independent features and global modality-shared features, respectively. Several simplified Transformer encoder layers form the Transformer backbone network, which is more suitable for real-time object tracking. Besides, we found that intermodality correlation is an important factor for modality interactions and mutual exploitation. Therefore, we propose a JMFL module, which uses cross-attention to capture the dependencies of cross-modal and enhance multimodal fusion by bidirectional guidance of multimodal information. The proposed method is fully experimented on two large benchmark datasets and compared with some current well-performing methods. The experimental results show that the proposed method performs well in terms of tracking accuracy and speed.
| 源语言 | 英语 |
|---|---|
| 期刊论文编号 | 6003805 |
| 期刊 | IEEE Geoscience and Remote Sensing Letters |
| 卷 | 20 |
| DOI | |
| 出版状态 | 已出版 - 2023 |
学术指纹
探究 'Visible and Infrared Object Tracking via Convolution-Transformer Network With Joint Multimodal Feature Learning' 的科研主题。它们共同构成独一无二的学术指纹。引用此
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver