Abstract
With the emergence of vision-language pre-trained models(VLM) like CLIP and ALIGN, transferable representations can be adapted to a wide range of downstream tasks via prompt tunning. While existing learnable prompt tuning approaches like CoOp and VPT have proved effective across varied tasks, their reliance on fixed parameters limits adaptability to new tasks without additional offline retraining. To adapt the pre-trained model to various tasks without the necessity for iterative offline training, we seek a task-driven prompt generator to map the tasks into prompts with a single feed-word process. In this paper, we propose a novel Generative Prompt (GPrompt) for vision-language pre-trained models, which is produced through Task-driven Prompt Generator (TDPG) driven by flexible and accurate task descriptors extracted from downstream samples. A flexible and simple inference pipeline is designed for efficient adaptation of our proposed method. Our method is able to surpass backpropagating-free prompt-based methods by +9.3%. Additionally, combined with existing training-free metric-based methods such as Tip-Adapter further improves the performance. Extensive experimental results on 11 datasets demonstrated the effectiveness of our proposed methods.
| Original language | English |
|---|---|
| Journal | IEEE Transactions on Multimedia |
| DOIs | |
| State | Accepted/In press - 2026 |
Keywords
- Few-shot Learning
- Image Classification
- Meta Learning
- Prompt Tunning
- Transfer Learning
- Vision-Language Model
Fingerprint
Dive into the research topics of 'Task-Driven Generative Prompt Learning for Vision-Language Pretrained Model'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver