Abstract
Recently, there has been growing interest in fine-tuning Vision-Language Models (VLMs) using extensive amounts of unlabeled data. Recent advances in this area can be broadly categorized into two approaches: self-training-based classifier tuning and feature tuning. Classifier tuning focuses on learning a customized classifier for the target domain while keeping the features invariant. In contrast, feature tuning aims to adjust the feature distribution of the target domain via visual prompt tuning to better align with a fixed classifier. We argue that due to imperfections in features and classifier, both approaches are essential for improving overall classification performance. To this end, we propose Fusing Classifier and Visual Prompt Tuning (FCVPT) for fine-tuning VLMs with unlabeled data. The core of FCVPT lies in the fusion of classifier tuning and visual prompt tuning during the fine-tuning process, leading to mutual enhancement. Additionally, we introduce a Neighborhood Consistency Constraint (NCC) for robust self-training. NCC exploits similarities between training examples in the feature space, encouraging predictions for each example to closely align with those of its nearest neighbors. Empowered by FCVPT and NCC, our method, dubbed FCVPT-NCC, achieves substantial performance enhancements compared to baselines across 14 datasets, showing an absolute improvement of up to 21.2% (with an average of 3.8%) to the best competitor. Code is available at: FCVPT-NCC.
| Original language | English |
|---|---|
| Article number | 109082 |
| Journal | Neural Networks |
| Volume | 203 |
| DOIs | |
| State | Published - Nov 2026 |
Keywords
- Clustering
- Prompt tuning
- Unsupervised learning
- Vision-language models
Fingerprint
Dive into the research topics of 'Unsupervised fine-tuning of vision-language models by fusing classifier tuning and visual prompt tuning'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver