TY - JOUR
T1 - Unsupervised fine-tuning of vision-language models by fusing classifier tuning and visual prompt tuning
AU - Chen, Wenyang
AU - Hu, Zhanxuan
AU - Tai, Yonghang
AU - Nie, Feiping
N1 - Publisher Copyright:
© 2026 Elsevier Ltd. All rights are reserved, including those for text and data mining, AI training, and similar technologies.
PY - 2026/11
Y1 - 2026/11
N2 - Recently, there has been growing interest in fine-tuning Vision-Language Models (VLMs) using extensive amounts of unlabeled data. Recent advances in this area can be broadly categorized into two approaches: self-training-based classifier tuning and feature tuning. Classifier tuning focuses on learning a customized classifier for the target domain while keeping the features invariant. In contrast, feature tuning aims to adjust the feature distribution of the target domain via visual prompt tuning to better align with a fixed classifier. We argue that due to imperfections in features and classifier, both approaches are essential for improving overall classification performance. To this end, we propose Fusing Classifier and Visual Prompt Tuning (FCVPT) for fine-tuning VLMs with unlabeled data. The core of FCVPT lies in the fusion of classifier tuning and visual prompt tuning during the fine-tuning process, leading to mutual enhancement. Additionally, we introduce a Neighborhood Consistency Constraint (NCC) for robust self-training. NCC exploits similarities between training examples in the feature space, encouraging predictions for each example to closely align with those of its nearest neighbors. Empowered by FCVPT and NCC, our method, dubbed FCVPT-NCC, achieves substantial performance enhancements compared to baselines across 14 datasets, showing an absolute improvement of up to 21.2% (with an average of 3.8%) to the best competitor. Code is available at: FCVPT-NCC.
AB - Recently, there has been growing interest in fine-tuning Vision-Language Models (VLMs) using extensive amounts of unlabeled data. Recent advances in this area can be broadly categorized into two approaches: self-training-based classifier tuning and feature tuning. Classifier tuning focuses on learning a customized classifier for the target domain while keeping the features invariant. In contrast, feature tuning aims to adjust the feature distribution of the target domain via visual prompt tuning to better align with a fixed classifier. We argue that due to imperfections in features and classifier, both approaches are essential for improving overall classification performance. To this end, we propose Fusing Classifier and Visual Prompt Tuning (FCVPT) for fine-tuning VLMs with unlabeled data. The core of FCVPT lies in the fusion of classifier tuning and visual prompt tuning during the fine-tuning process, leading to mutual enhancement. Additionally, we introduce a Neighborhood Consistency Constraint (NCC) for robust self-training. NCC exploits similarities between training examples in the feature space, encouraging predictions for each example to closely align with those of its nearest neighbors. Empowered by FCVPT and NCC, our method, dubbed FCVPT-NCC, achieves substantial performance enhancements compared to baselines across 14 datasets, showing an absolute improvement of up to 21.2% (with an average of 3.8%) to the best competitor. Code is available at: FCVPT-NCC.
KW - Clustering
KW - Prompt tuning
KW - Unsupervised learning
KW - Vision-language models
UR - https://www.scopus.com/pages/publications/105038893697
U2 - 10.1016/j.neunet.2026.109082
DO - 10.1016/j.neunet.2026.109082
M3 - 文章
AN - SCOPUS:105038893697
SN - 0893-6080
VL - 203
JO - Neural Networks
JF - Neural Networks
M1 - 109082
ER -