跳到主要导航 跳到搜索 跳到主要内容

A Survey of Medical Vision-and-Language Applications and Their Techniques

  • Qi Chen
  • , Ruoshan Zhao
  • , Sinuo Wang
  • , Vu Minh Hieu Phan
  • , Anton van den Hengel
  • , Johan Verjans
  • , Zhibin Liao
  • , Minh Son To
  • , Yong Xia
  • , Jian Chen
  • , Yutong Xie
  • , Qi Wu
  • University of Adelaide
  • South China University of Technology
  • Flinders University

科研成果: 期刊稿件文章同行评审

摘要

Medical vision-and-language models (MVLMs) have attracted substantial interest due to their capability to offer a natural language interface for interpreting complex medical data. Their applications are versatile and have the potential to improve diagnostic accuracy and decision-making for individual patients while also contributing to enhanced public health monitoring, disease surveillance, and policy-making through more efficient analysis of large data sets. MVLMS integrate natural language processing with medical images to enable a more comprehensive and contextual understanding of medical images alongside their corresponding textual information. Unlike general vision-and-language models trained on diverse, non-specialized datasets, MVLMs are purpose-built for the medical domain, automatically extracting and interpreting critical information from medical images and textual reports to support clinical decision-making. Popular clinical applications of MVLMs include automated medical report generation, medical visual question answering, medical multimodal segmentation, diagnosis and prognosis and medical image-text retrieval. Here, we provide a comprehensive overview of MVLMs and the various medical tasks to which they have been applied. We conduct a detailed analysis of various vision-and-language model architectures, focusing on their distinct strategies for cross-modal integration/exploitation of medical visual and textual features. We also examine the datasets used for these tasks and compare the performance of different models based on standardized evaluation metrics. Furthermore, we highlight potential challenges and summarize future research trends and directions. The full collection of papers and codes is available at: https://github.com/YtongXie/Medical-Vision-and-Language-Tasks-and-Methodologies-A-Survey.

源语言英语
期刊论文编号381
期刊International Journal of Computer Vision
134
8
DOI
出版状态已出版 - 8月 2026

联合国可持续发展目标

此成果有助于实现下列可持续发展目标:

  1. 可持续发展目标 3 - 良好健康与福祉
    可持续发展目标 3 良好健康与福祉

学术指纹

探究 'A Survey of Medical Vision-and-Language Applications and Their Techniques' 的科研主题。它们共同构成独一无二的学术指纹。

引用此