Skip to main navigation Skip to search Skip to main content

VSCode-v2: Dynamic Prompt Learning for General Visual Salient and Camouflaged Object Detection With Two-Stage Optimization

  • Northwestern Polytechnical University Xian
  • Nankai University
  • Nankai International Advanced Research Institute
  • Mohamed Bin Zayed University of Artificial Intelligence
  • Chongqing University of Posts and Telecommunications

Research output: Contribution to journalArticlepeer-review

1 Scopus citations

Abstract

Salient object detection (SOD) and camouflaged object detection (COD) are related but distinct binary mapping tasks, each involving multiple modalities that share commonalities while maintaining unique characteristics. Existing approaches often rely on complex, task-specific architectures, leading to redundancy and limited generalization. Our previous work, VSCode, introduced a generalist model that effectively handles four SOD tasks and two COD tasks. VSCode leveraged VST as its foundation model and incorporated 2D prompts within an encoder-decoder framework to capture domain and task-specific knowledge, utilizing a prompt discrimination loss to optimize the model. Building upon the proven effectiveness of our previous work VSCode, we identify opportunities to further strengthen generalization capabilities through focused modifications in model design and optimization strategy. To unlock this potential, we propose VSCode-v2, an extension that introduces a Mixture of Prompt Experts (MoPE) layer to generate adaptive prompts. We also redesign the training process into a two-stage approach: first learning shared features across tasks, then capturing specific characteristics. To preserve knowledge during this process, we incorporate distillation from our conference version model. Furthermore, we propose a contrastive learning mechanism with data augmentation to strengthen the relationships between prompts and feature representations. VSCode-v2 demonstrates balanced performance improvements across six SOD and COD tasks. Moreover, VSCode-v2 effectively handles various multimodal inputs and exhibits zero-shot generalization capability to novel tasks, such as RGB-D Video SOD.

Original languageEnglish
Pages (from-to)3137-3153
Number of pages17
JournalIEEE Transactions on Pattern Analysis and Machine Intelligence
Volume48
Issue number3
DOIs
StatePublished - 2026

Keywords

  • Saliency object detection
  • camouflaged object detection
  • multi-task learning
  • prompt

Fingerprint

Dive into the research topics of 'VSCode-v2: Dynamic Prompt Learning for General Visual Salient and Camouflaged Object Detection With Two-Stage Optimization'. Together they form a unique fingerprint.

Cite this