跳到主要导航 跳到搜索 跳到主要内容

SNP-S3: Shared Network Pre-Training and Significant Semantic Strengthening for Various Video-Text Tasks

  • Xingning Dong
  • , Qingpei Guo
  • , Tian Gan
  • , Qing Wang
  • , Jianlong Wu
  • , Xiangyuan Ren
  • , Yuan Cheng
  • , Wei Chu
  • Shandong University
  • Ant Group
  • Harbin Institute of Technology (Shenzhen)
  • Fudan University

科研成果: 期刊稿件文章同行评审

5 引用 (Scopus)

摘要

We present a framework for learning cross-modal video representations by directly pre-training on raw data to facilitate various downstream video-text tasks. Our main contributions lie in the pre-training framework and proxy tasks. First, based on the shortcomings of two mainstream pixel-level pre-training architectures (limited applications or less efficient), we propose Shared Network Pre-training (SNP). By employing one shared BERT-type network to refine textual and cross-modal features simultaneously, SNP is lightweight and could support various downstream applications. Second, based on the intuition that people always pay attention to several “significant words” when understanding a sentence, we propose the Significant Semantic Strengthening (S3) strategy, which includes a novel masking and matching proxy task to promote the pre-training performance. Experiments conducted on three downstream video-text tasks and six datasets demonstrate that, we establish a new state-of-the-art in pixel-level video-text pre-training; we also achieve a satisfactory balance between the pre-training efficiency and the fine-tuning performance. The codebase and pre-trained models are available at https://github.com/dongxingning/SNPS3.

源语言英语
页(从-至)2525-2535
页数11
期刊IEEE Transactions on Circuits and Systems for Video Technology
34
4
DOI
出版状态已出版 - 1 4月 2024
已对外发布

学术指纹

探究 'SNP-S3: Shared Network Pre-Training and Significant Semantic Strengthening for Various Video-Text Tasks' 的科研主题。它们共同构成独一无二的学术指纹。

引用此