跳到主要导航 跳到搜索 跳到主要内容

Semi-Supervised VQA Multi-Modal Explanation via Self-Critical Learning

  • Wei Suo
  • , Ji Ma
  • , Mengyang Sun
  • , Hanwang Zhang
  • , Peng Wang
  • , Yanning Zhang
  • , Qi Wu
  • Northwestern Polytechnical University Xian
  • Nanyang Technological University
  • University of Adelaide

科研成果: 期刊稿件文章同行评审

摘要

VQA explanation task aims to explain the decision-making process of VQA models in a way that is easily understandable to humans. Existing methods mostly use visual location or natural language explanation approaches to generate corresponding rationales. Although significant progress has been made, these frameworks are bottlenecked by the following challenges: 1) Uni-modal paradigm inevitably leads to semantic ambiguity of explanations. 2) The reasoning process cannot be faithfully responded to and suffers from logical inconsistency. 3) Human-annotated explanations are expensive and time-consuming to collect. In this paper, we introduce a new Semi-supervised VQA Multi-modal Explanation (SME) method via self-critical learning, which addresses the above challenges by leveraging both visual and textual explanations to comprehensively reveal the inference process of the model. Meanwhile, in order to improve the logical consistency between answers and rationales, we design a novel self-critical strategy to evaluate candidate explanations based on answer reward scores. More importantly, our method can benefit from a tremendous amount of samples without human-annotated explanations with semi-supervised learning. Extensive automatic measures and human evaluations all show the effectiveness of our method. Finally, the framework achieves a new state-of-the-art performance on the three VQA explanation datasets.

源语言英语
页(从-至)8361-8377
页数17
期刊IEEE Transactions on Pattern Analysis and Machine Intelligence
48
7
DOI
出版状态已出版 - 1 7月 2026

学术指纹

探究 'Semi-Supervised VQA Multi-Modal Explanation via Self-Critical Learning' 的科研主题。它们共同构成独一无二的学术指纹。

引用此