Abstract
With the increasing prevalence of surveillance devices, automatic detection of anomalies in videos has gained significant importance. However, traditional methods tend to utilize only visual information while often neglecting the substantial expressive potential of textual data. To address this limitation, this study proposes an innovative framework that improves the detection performance by simultaneously leveraging both coarse-grained label text and fine-grained description text to enhance vision information. Specifically, our framework consists of three core components: the temporal enhanced module to capture local and global temporal dependencies of the video, the label enhanced module to establish high-level semantic associations using label text, and the description enhanced module to achieve fine-grained contextual alignment using description text. Experimental results on the UCF-Crime and XD-Violence datasets demonstrate that our approach achieves excellent performance, thereby validating its efficacy in accurately detecting anomalous events within complex scenarios.
| Original language | English |
|---|---|
| Title of host publication | 2025 6th International Conference on Computer Vision, Image and Deep Learning, CVIDL 2025 |
| Publisher | Institute of Electrical and Electronics Engineers Inc. |
| Pages | 764-767 |
| Number of pages | 4 |
| ISBN (Electronic) | 9798331523244 |
| DOIs | |
| State | Published - 2025 |
| Event | 6th International Conference on Computer Vision, Image and Deep Learning, CVIDL 2025 - Ningbo, China Duration: 23 May 2025 → 25 May 2025 |
Publication series
| Name | 2025 6th International Conference on Computer Vision, Image and Deep Learning, CVIDL 2025 |
|---|
Conference
| Conference | 6th International Conference on Computer Vision, Image and Deep Learning, CVIDL 2025 |
|---|---|
| Country/Territory | China |
| City | Ningbo |
| Period | 23/05/25 → 25/05/25 |
UN SDGs
This output contributes to the following UN Sustainable Development Goals (SDGs)
-
SDG 16 Peace, Justice and Strong Institutions
Keywords
- Anomaly Detection
- Multimodal framework
- Weak Supervision
Fingerprint
Dive into the research topics of 'Multi-Granularity Text Enhanced Framework for Weakly Supervised Video Anomaly Detection Via Cross-Modal Context Alignment'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver