Abstract
With the extensive deployment of surveillance cameras, Weakly Supervised Video Anomaly Detection (WSVAD) has attracted increasing attention in many fields. It significantly reduces the labeling cost by relying only on video-level labels for training, and shows important significance in practical applications. However, existing methods often depend on unimodal visual information, neglecting the rich semantic information embedded in video description text. To address this limitation, this paper proposes a novel framework: Generative Description Boosted Weakly Supervised Video Anomaly Detection (DBVAD). DBVAD leverages large vision language models as the knowledge engine to generate video descriptions, which are then utilized as semantic supervision signals to optimize visual features. The proposed DBVAD comprises several key components. First, the key event selection strategy is used to accurately select key frames from videos for subsequent description generation. Second, the temporal modeling module captures the multi-scale temporal dependencies within videos. Lastly, the semantic focus prompt calibrates visual representations using label texts, while the description boosted module achieves fine alignment between visual features and generated description text through contrastive learning, thereby enhancing the model’s semantic understanding of abnormal events. Experimental results indicate that DBVAD achieves superior performance on the large-scale UCF-Crime and XD-Violence datasets, thereby validating its effectiveness.
| Original language | English |
|---|---|
| Title of host publication | Pattern Recognition and Computer Vision - 8th Chinese Conference, PRCV 2025, Proceedings |
| Editors | Josef Kittler, Hongkai Xiong, Weiyao Lin, Jian Yang, Xilin Chen, Jiwen Lu, Jingyi Yu, Weishi Zheng |
| Publisher | Springer Science and Business Media Deutschland GmbH |
| Pages | 358-372 |
| Number of pages | 15 |
| ISBN (Print) | 9789819555666 |
| DOIs | |
| State | Published - 2026 |
| Event | 8th Chinese Conference on Pattern Recognition and Computer Vision, PRCV 2025 - Shanghai, China Duration: 15 Oct 2025 → 18 Oct 2025 |
Publication series
| Name | Lecture Notes in Computer Science |
|---|---|
| Volume | 16276 LNCS |
| ISSN (Print) | 0302-9743 |
| ISSN (Electronic) | 1611-3349 |
Conference
| Conference | 8th Chinese Conference on Pattern Recognition and Computer Vision, PRCV 2025 |
|---|---|
| Country/Territory | China |
| City | Shanghai |
| Period | 15/10/25 → 18/10/25 |
UN SDGs
This output contributes to the following UN Sustainable Development Goals (SDGs)
-
SDG 16 Peace, Justice and Strong Institutions
Keywords
- Anomaly Detection
- Multimodal Framework
- Weak Supervision
Fingerprint
Dive into the research topics of 'Boosting Weakly Supervised Video Anomaly Detection with Generative Description'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver