Skip to main navigation Skip to search Skip to main content

Boosting Weakly Supervised Video Anomaly Detection with Generative Description

  • Chenlin Meng
  • , Zhaoyong Mao
  • , Chi Zhang
  • , Kai Jiang
  • , Junge Shen
  • Northwestern Polytechnical University Xian

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

Abstract

With the extensive deployment of surveillance cameras, Weakly Supervised Video Anomaly Detection (WSVAD) has attracted increasing attention in many fields. It significantly reduces the labeling cost by relying only on video-level labels for training, and shows important significance in practical applications. However, existing methods often depend on unimodal visual information, neglecting the rich semantic information embedded in video description text. To address this limitation, this paper proposes a novel framework: Generative Description Boosted Weakly Supervised Video Anomaly Detection (DBVAD). DBVAD leverages large vision language models as the knowledge engine to generate video descriptions, which are then utilized as semantic supervision signals to optimize visual features. The proposed DBVAD comprises several key components. First, the key event selection strategy is used to accurately select key frames from videos for subsequent description generation. Second, the temporal modeling module captures the multi-scale temporal dependencies within videos. Lastly, the semantic focus prompt calibrates visual representations using label texts, while the description boosted module achieves fine alignment between visual features and generated description text through contrastive learning, thereby enhancing the model’s semantic understanding of abnormal events. Experimental results indicate that DBVAD achieves superior performance on the large-scale UCF-Crime and XD-Violence datasets, thereby validating its effectiveness.

Original languageEnglish
Title of host publicationPattern Recognition and Computer Vision - 8th Chinese Conference, PRCV 2025, Proceedings
EditorsJosef Kittler, Hongkai Xiong, Weiyao Lin, Jian Yang, Xilin Chen, Jiwen Lu, Jingyi Yu, Weishi Zheng
PublisherSpringer Science and Business Media Deutschland GmbH
Pages358-372
Number of pages15
ISBN (Print)9789819555666
DOIs
StatePublished - 2026
Event8th Chinese Conference on Pattern Recognition and Computer Vision, PRCV 2025 - Shanghai, China
Duration: 15 Oct 202518 Oct 2025

Publication series

NameLecture Notes in Computer Science
Volume16276 LNCS
ISSN (Print)0302-9743
ISSN (Electronic)1611-3349

Conference

Conference8th Chinese Conference on Pattern Recognition and Computer Vision, PRCV 2025
Country/TerritoryChina
CityShanghai
Period15/10/2518/10/25

UN SDGs

This output contributes to the following UN Sustainable Development Goals (SDGs)

  1. SDG 16 - Peace, Justice and Strong Institutions
    SDG 16 Peace, Justice and Strong Institutions

Keywords

  • Anomaly Detection
  • Multimodal Framework
  • Weak Supervision

Fingerprint

Dive into the research topics of 'Boosting Weakly Supervised Video Anomaly Detection with Generative Description'. Together they form a unique fingerprint.

Cite this