跳到主要导航 跳到搜索 跳到主要内容

Dual-Branch Multimodal Graph Learning for Contextual Semantic Enhanced Video Captioning

  • Northwestern Polytechnical University Xian

科研成果: 期刊稿件文章同行评审

摘要

Video captioning, aiming at describing the content of given videos in natural language, has become an interesting research area and has gained broad attention. Although significant progress has been made, current graph-based caption methods often focus solely on modeling visual relationships, neglecting the interaction between visual and linguistic information. And they also fail to capture the rich semantic-level correlation present in video and text, hindering a deeper understanding of the complex relationship within multimodal content. In this paper, we introduce a dual-branch multimodal graph learning method for contextual semantic enhanced video captioning, where apparent structure-level and latent semantic-level multimodal graphs are constructed and jointly utilized for the caption generation. The apparent structure-level multimodal graph emphasizes the interactions between video and text, providing a comprehensive representation of multimodal content at the apparent level. Meanwhile, the latent semantic-level multimodal graph focuses on deeper semantic relationships between visual and textual elements, capturing their underlying associations for a more refined understanding of implicit meanings at the latent semantic level. Moreover, a cross-graph representation fusion mechanism is designed to effectively fuse and extract consistent representations from both structure-level and semantic-level multimodal graphs. Experimental results on several popular datasets demonstrate the effectiveness of the proposed method, which produces higher-quality caption results.

源语言英语
期刊IEEE Transactions on Multimedia
DOI
出版状态已接受/待刊 - 2026

指纹

探究 'Dual-Branch Multimodal Graph Learning for Contextual Semantic Enhanced Video Captioning' 的科研主题。它们共同构成独一无二的指纹。

引用此