摘要
- Weakly supervised video anomaly detection (WS-VAD) is challenging because it relies on video-level binary annotations to make frame-level predictions. Existing methods often convert WSVAD into a multiple instance learning (MIL) task, focusing on isolated segments that contribute most to the classification while neglecting the temporal context and detailed feature distinctions. In this paper, we propose a contrastive clustering strategy that enhances the representation of normal and abnormal features. Specifically, we treat the clustering center features and their corresponding categories as positive sample pairs, while features from different categories are treated as negative samples. This approach enables the network to better explore the distinction between normal and abnormal features. Furthermore, we address the bias in pre-trained models, where I3D pre-training features tend to overfit to normal videos and CLIP features exhibit a bias towards abnormal videos. To mitigate this, we introduce a simple early fusion method that combines pre-trained features to eliminate bias and obtain more comprehensive spatio-temporal representations. Extensive experiments on the UCF-Crime and XD-Violence datasets demonstrate the effectiveness of our approach, achieving state-of-the-art performance.
| 源语言 | 英语 |
|---|---|
| 期刊 | Proceedings of the International Joint Conference on Neural Networks |
| DOI | |
| 出版状态 | 已出版 - 2025 |
| 活动 | 2025 International Joint Conference on Neural Networks, IJCNN 2025 - Rome, 意大利 期限: 30 6月 2025 → 5 7月 2025 |
联合国可持续发展目标
此成果有助于实现下列可持续发展目标:
-
可持续发展目标 16 和平、正义和强大机构
指纹
探究 'Weakly Supervised Video Anomaly Detection Via Contrastive Clustering' 的科研主题。它们共同构成独一无二的指纹。引用此
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver