跳到主要导航 跳到搜索 跳到主要内容

Aligning local features from multi-view (ALFM): A hybrid self-Supervised framework for object detection via contextual distillation and global representation learning

  • Zhenyu Fang
  • , Zhuowei Wang
  • , Jinchang Ren
  • , Jiangbin Zheng
  • , Rongjun Chen
  • , Huimin Zhao
  • Northwestern Polytechnical University Xian
  • Guangdong Polytechnic Normal University
  • Robert Gordon University

科研成果: 期刊稿件文章同行评审

2 引用 (Scopus)

摘要

Self-supervised learning learns generalized representations using unlabeled data for downstream tasks, where many of them are optimized with a pretext task derived from multi-view augmented images, based on the assumption that the majority of the foreground objects in the source dataset and the background features are redundant. However, for object detection datasets, background features are essential for accurately detecting objects. In this paper, a detection-specific self-supervised method is proposed for aligning local features from multi-view images (ALFM). The proposed ALFM consists of two learning branches: global minimal sufficient representation (GMSR) and contextual distillation on local patches (CDLP). The GMSR loss globally learns sufficient feature representations with minimal redundant information, enabling the network to maintain generalization when foreground categories are not determined. This is achieved by maximizing the similarity between the embeddings of two views and increasing the differential entropy of the embeddings from each view. The CDLP loss is proposed to enhance local feature representations while reducing redundant information caused by the gap between the pretext task and the detection task. This is achieved by learning to predict ”soft-labels” with rich contextual information. Taking COCO as the pretraining dataset, results from various detection benchmarks validate the efficacy of the proposed ALFM, achieving similar mAP as ImageNet-pretrained models while using only 10 % of the training samples.

源语言英语
期刊论文编号114671
期刊Knowledge-Based Systems
330
DOI
出版状态已出版 - 25 11月 2025

学术指纹

探究 'Aligning local features from multi-view (ALFM): A hybrid self-Supervised framework for object detection via contextual distillation and global representation learning' 的科研主题。它们共同构成独一无二的学术指纹。

引用此