跳到主要导航 跳到搜索 跳到主要内容

Task-Driven Generative Prompt Learning for Vision-Language Pretrained Model

  • Qirui Wu
  • , Shizhou Zhang
  • , De Cheng
  • , Yinghui Xing
  • , Guoqiang Liang
  • , Kairui Dang
  • , Yanning Zhang
  • Northwestern Polytechnical University Xian
  • Xidian University

科研成果: 期刊稿件文章同行评审

摘要

With the emergence of vision-language pre-trained models(VLM) like CLIP and ALIGN, transferable representations can be adapted to a wide range of downstream tasks via prompt tunning. While existing learnable prompt tuning approaches like CoOp and VPT have proved effective across varied tasks, their reliance on fixed parameters limits adaptability to new tasks without additional offline retraining. To adapt the pre-trained model to various tasks without the necessity for iterative offline training, we seek a task-driven prompt generator to map the tasks into prompts with a single feed-word process. In this paper, we propose a novel Generative Prompt (GPrompt) for vision-language pre-trained models, which is produced through Task-driven Prompt Generator (TDPG) driven by flexible and accurate task descriptors extracted from downstream samples. A flexible and simple inference pipeline is designed for efficient adaptation of our proposed method. Our method is able to surpass backpropagating-free prompt-based methods by +9.3%. Additionally, combined with existing training-free metric-based methods such as Tip-Adapter further improves the performance. Extensive experimental results on 11 datasets demonstrated the effectiveness of our proposed methods.

源语言英语
期刊IEEE Transactions on Multimedia
DOI
出版状态已接受/待刊 - 2026

指纹

探究 'Task-Driven Generative Prompt Learning for Vision-Language Pretrained Model' 的科研主题。它们共同构成独一无二的指纹。

引用此