跳到主要导航 跳到搜索 跳到主要内容

Unsupervised fine-tuning of vision-language models by fusing classifier tuning and visual prompt tuning

  • Yunnan Normal University

科研成果: 期刊稿件文章同行评审

摘要

Recently, there has been growing interest in fine-tuning Vision-Language Models (VLMs) using extensive amounts of unlabeled data. Recent advances in this area can be broadly categorized into two approaches: self-training-based classifier tuning and feature tuning. Classifier tuning focuses on learning a customized classifier for the target domain while keeping the features invariant. In contrast, feature tuning aims to adjust the feature distribution of the target domain via visual prompt tuning to better align with a fixed classifier. We argue that due to imperfections in features and classifier, both approaches are essential for improving overall classification performance. To this end, we propose Fusing Classifier and Visual Prompt Tuning (FCVPT) for fine-tuning VLMs with unlabeled data. The core of FCVPT lies in the fusion of classifier tuning and visual prompt tuning during the fine-tuning process, leading to mutual enhancement. Additionally, we introduce a Neighborhood Consistency Constraint (NCC) for robust self-training. NCC exploits similarities between training examples in the feature space, encouraging predictions for each example to closely align with those of its nearest neighbors. Empowered by FCVPT and NCC, our method, dubbed FCVPT-NCC, achieves substantial performance enhancements compared to baselines across 14 datasets, showing an absolute improvement of up to 21.2% (with an average of 3.8%) to the best competitor. Code is available at: FCVPT-NCC.

源语言英语
文章编号109082
期刊Neural Networks
203
DOI
出版状态已出版 - 11月 2026

指纹

探究 'Unsupervised fine-tuning of vision-language models by fusing classifier tuning and visual prompt tuning' 的科研主题。它们共同构成独一无二的指纹。

引用此