摘要
With the emergence of vision-language pre-trained models(VLM) like CLIP and ALIGN, transferable representations can be adapted to a wide range of downstream tasks via prompt tunning. While existing learnable prompt tuning approaches like CoOp and VPT have proved effective across varied tasks, their reliance on fixed parameters limits adaptability to new tasks without additional offline retraining. To adapt the pre-trained model to various tasks without the necessity for iterative offline training, we seek a task-driven prompt generator to map the tasks into prompts with a single feed-word process. In this paper, we propose a novel Generative Prompt (GPrompt) for vision-language pre-trained models, which is produced through Task-driven Prompt Generator (TDPG) driven by flexible and accurate task descriptors extracted from downstream samples. A flexible and simple inference pipeline is designed for efficient adaptation of our proposed method. Our method is able to surpass backpropagating-free prompt-based methods by +9.3%. Additionally, combined with existing training-free metric-based methods such as Tip-Adapter further improves the performance. Extensive experimental results on 11 datasets demonstrated the effectiveness of our proposed methods.
| 源语言 | 英语 |
|---|---|
| 期刊 | IEEE Transactions on Multimedia |
| DOI | |
| 出版状态 | 已接受/待刊 - 2026 |
指纹
探究 'Task-Driven Generative Prompt Learning for Vision-Language Pretrained Model' 的科研主题。它们共同构成独一无二的指纹。引用此
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver