跳到主要导航 跳到搜索 跳到主要内容

Single-frame supervision for temporal video anomaly grounding

  • Yuzhou Long
  • , Peng Wu
  • , Yuting Yan
  • , Guansong Pang
  • , Peng Wang
  • , Yanning Zhang
  • Northwestern Polytechnical University Xian
  • Singapore Management University

科研成果: 期刊稿件文章同行评审

摘要

Conventional video anomaly detection approaches struggle with increasingly sophisticated fine-grained analysis requirements in real-world applications, establishing Temporal Video Anomaly Grounding (TVAG) as one of the pivotal research frontiers in advanced anomaly video comprehension systems. Targeting the scarcity of precise temporal annotations, this work develops a single-frame supervision-based framework, Glance-guided Cross-modal Proposal Generation (GCPG), which offers competitive grounding performance, surpassing some fully supervised methods under specific metrics, while substantially reducing annotation costs. The framework consists of a Cross-Modal Collaborative Pseudo-Glance Localization module (PGL) and a Glance-Guided Gaussian Proposal Optimization module (GPO). PGL employs a semantic-aware dual-branch mechanism that jointly performs cross-modal feature fusion classification and textual semantic verification to generate reliable pseudo-frame supervision, forming the foundation for cross-modal alignment learning. GPO enhances proposal quality by reconstructing Gaussian mask composition weights based on glance-keyword alignment and distributional consistency. Comprehensive experiments and ablation analyses on two challenging TVAG benchmarks validate the efficacy of our single-frame supervised approach.

源语言英语
文章编号132346
期刊Neurocomputing
668
DOI
出版状态已出版 - 1 3月 2026

指纹

探究 'Single-frame supervision for temporal video anomaly grounding' 的科研主题。它们共同构成独一无二的指纹。

引用此