TY - JOUR
T1 - Aligning local features from multi-view (ALFM)
T2 - A hybrid self-Supervised framework for object detection via contextual distillation and global representation learning
AU - Fang, Zhenyu
AU - Wang, Zhuowei
AU - Ren, Jinchang
AU - Zheng, Jiangbin
AU - Chen, Rongjun
AU - Zhao, Huimin
N1 - Publisher Copyright:
© 2025
PY - 2025/11/25
Y1 - 2025/11/25
N2 - Self-supervised learning learns generalized representations using unlabeled data for downstream tasks, where many of them are optimized with a pretext task derived from multi-view augmented images, based on the assumption that the majority of the foreground objects in the source dataset and the background features are redundant. However, for object detection datasets, background features are essential for accurately detecting objects. In this paper, a detection-specific self-supervised method is proposed for aligning local features from multi-view images (ALFM). The proposed ALFM consists of two learning branches: global minimal sufficient representation (GMSR) and contextual distillation on local patches (CDLP). The GMSR loss globally learns sufficient feature representations with minimal redundant information, enabling the network to maintain generalization when foreground categories are not determined. This is achieved by maximizing the similarity between the embeddings of two views and increasing the differential entropy of the embeddings from each view. The CDLP loss is proposed to enhance local feature representations while reducing redundant information caused by the gap between the pretext task and the detection task. This is achieved by learning to predict ”soft-labels” with rich contextual information. Taking COCO as the pretraining dataset, results from various detection benchmarks validate the efficacy of the proposed ALFM, achieving similar mAP as ImageNet-pretrained models while using only 10 % of the training samples.
AB - Self-supervised learning learns generalized representations using unlabeled data for downstream tasks, where many of them are optimized with a pretext task derived from multi-view augmented images, based on the assumption that the majority of the foreground objects in the source dataset and the background features are redundant. However, for object detection datasets, background features are essential for accurately detecting objects. In this paper, a detection-specific self-supervised method is proposed for aligning local features from multi-view images (ALFM). The proposed ALFM consists of two learning branches: global minimal sufficient representation (GMSR) and contextual distillation on local patches (CDLP). The GMSR loss globally learns sufficient feature representations with minimal redundant information, enabling the network to maintain generalization when foreground categories are not determined. This is achieved by maximizing the similarity between the embeddings of two views and increasing the differential entropy of the embeddings from each view. The CDLP loss is proposed to enhance local feature representations while reducing redundant information caused by the gap between the pretext task and the detection task. This is achieved by learning to predict ”soft-labels” with rich contextual information. Taking COCO as the pretraining dataset, results from various detection benchmarks validate the efficacy of the proposed ALFM, achieving similar mAP as ImageNet-pretrained models while using only 10 % of the training samples.
KW - Contextual distillation
KW - Global minimal sufficient representation
KW - Object detection
KW - Self-supervised learning
UR - https://www.scopus.com/pages/publications/105020266528
U2 - 10.1016/j.knosys.2025.114671
DO - 10.1016/j.knosys.2025.114671
M3 - 文章
AN - SCOPUS:105020266528
SN - 0950-7051
VL - 330
JO - Knowledge-Based Systems
JF - Knowledge-Based Systems
M1 - 114671
ER -