跳到主要导航 跳到搜索 跳到主要内容

Adaptive Medical Topic Learning for Enhanced Fine-Grained Cross-Modal Alignment in Medical Report Generation

  • Xin Mei
  • , Libin Yang
  • , Dehong Gao
  • , Xiaoyan Cai
  • , Junwei Han
  • , Tianming Liu
  • Northwestern Polytechnical University Xian
  • University of Georgia

科研成果: 期刊稿件文章同行评审

11 引用 (Scopus)

摘要

Medical report generation refers to the automatic creation of accurate and coherent diagnostic reports for medical images. This task can alleviate the workload of radiologists, enhance the efficiency of disease diagnosis, and therefore holds significant value and challenges. Considering the feature differences between different modalities, existing methods primarily focus on facilitating medical report generation through cross-modal alignment of images and texts. However, since medical images are very similar to each other, it is difficult to tag obvious objects, making most methods limited to coarse-grained image-text global alignment. In this paper, we propose a medical report generation model based on adaptive topic learning and fine-grained cross-modal alignment, which aligns images and texts from medical topic perspective and token perspective. From the medical topic perspective, a global-local contrastive loss is introduced to adaptively learn efficient medical topic features, and medical topics are utilized to map images and texts to the same semantic space for fine-grained alignment. From the token perspective, a token prediction module is designed to enable the model to focus on important local information by predicting the key tokens contained in the report. Experimental results on the two public datasets (i.e. IU-Xray and MIMIC-CXR) demonstrate that our proposed model outperforms state-of-the-art baselines.

源语言英语
页(从-至)5050-5061
页数12
期刊IEEE Transactions on Multimedia
27
DOI
出版状态已出版 - 2025

学术指纹

探究 'Adaptive Medical Topic Learning for Enhanced Fine-Grained Cross-Modal Alignment in Medical Report Generation' 的科研主题。它们共同构成独一无二的学术指纹。

引用此