跳到主要导航 跳到搜索 跳到主要内容

Video summarization with a convolutional attentive adversarial network

  • Guoqiang Liang
  • , Yanbing Lv
  • , Shucheng Li
  • , Shizhou Zhang
  • , Yanning Zhang
  • Northwestern Polytechnical University Xian

科研成果: 期刊稿件文章同行评审

33 引用 (Scopus)

摘要

With the explosive growth of video data, video summarization, which attempts to seek the minimum subset of frames while still conveying the main story, has become one of the hottest topics. Nowadays, substantial achievements have been made by supervised learning techniques, especially after the emergence of deep learning. However, it is extremely expensive and difficult to construct a large-scale video summarization dataset through human annotation. To address this problem, we propose a convolutional attentive adversarial network (CAAN), whose key idea is to build a deep summarizer in an unsupervised way. Upon the generative adversarial network, our overall framework consists of a generator and a discriminator. The former predicts importance scores for all the frames of a video while the latter tries to distinguish the score-weighted frame features from original frame features. To capture the global and local temporal relationship of video frames, the generator employs a fully convolutional sequence network to build global representation of a video, and an attention-based network to predict normalized importance scores. To optimize the parameters, our objective function is composed of three loss functions, which can guide the frame-level importance score prediction collaboratively. To validate this proposed method, we have conducted extensive experiments on two public benchmarks SumMe and TVSum. The results show the superiority of our proposed method against other state-of-the-art unsupervised approaches. Our method even outperforms some published supervised approaches.

源语言英语
文章编号108840
期刊Pattern Recognition
131
DOI
出版状态已出版 - 11月 2022

指纹

探究 'Video summarization with a convolutional attentive adversarial network' 的科研主题。它们共同构成独一无二的指纹。

引用此