Skip to main navigation Skip to search Skip to main content

Multi-Granularity Text Enhanced Framework for Weakly Supervised Video Anomaly Detection Via Cross-Modal Context Alignment

  • Northwestern Polytechnical University Xian

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

1 Scopus citations

Abstract

With the increasing prevalence of surveillance devices, automatic detection of anomalies in videos has gained significant importance. However, traditional methods tend to utilize only visual information while often neglecting the substantial expressive potential of textual data. To address this limitation, this study proposes an innovative framework that improves the detection performance by simultaneously leveraging both coarse-grained label text and fine-grained description text to enhance vision information. Specifically, our framework consists of three core components: the temporal enhanced module to capture local and global temporal dependencies of the video, the label enhanced module to establish high-level semantic associations using label text, and the description enhanced module to achieve fine-grained contextual alignment using description text. Experimental results on the UCF-Crime and XD-Violence datasets demonstrate that our approach achieves excellent performance, thereby validating its efficacy in accurately detecting anomalous events within complex scenarios.

Original languageEnglish
Title of host publication2025 6th International Conference on Computer Vision, Image and Deep Learning, CVIDL 2025
PublisherInstitute of Electrical and Electronics Engineers Inc.
Pages764-767
Number of pages4
ISBN (Electronic)9798331523244
DOIs
StatePublished - 2025
Event6th International Conference on Computer Vision, Image and Deep Learning, CVIDL 2025 - Ningbo, China
Duration: 23 May 202525 May 2025

Publication series

Name2025 6th International Conference on Computer Vision, Image and Deep Learning, CVIDL 2025

Conference

Conference6th International Conference on Computer Vision, Image and Deep Learning, CVIDL 2025
Country/TerritoryChina
CityNingbo
Period23/05/2525/05/25

UN SDGs

This output contributes to the following UN Sustainable Development Goals (SDGs)

  1. SDG 16 - Peace, Justice and Strong Institutions
    SDG 16 Peace, Justice and Strong Institutions

Keywords

  • Anomaly Detection
  • Multimodal framework
  • Weak Supervision

Fingerprint

Dive into the research topics of 'Multi-Granularity Text Enhanced Framework for Weakly Supervised Video Anomaly Detection Via Cross-Modal Context Alignment'. Together they form a unique fingerprint.

Cite this