Skip to main navigation Skip to search Skip to main content

Task-Agnostic Pre-training and Task-Guided Fine-tuning for Versatile Diffusion Planner

  • Chenyou Fan
  • , Chenjia Bai
  • , Zhao Shan
  • , Haoran He
  • , Yang Zhang
  • , Zhen Wang
  • Northwestern Polytechnical University Xian
  • China Telecommunications
  • Tsinghua University
  • Hong Kong University of Science and Technology

Research output: Contribution to journalConference articlepeer-review

Abstract

Diffusion models have demonstrated their capabilities in modeling trajectories of multi-tasks. However, existing multi-task planners or policies typically rely on task-specific demonstrations via multi-task imitation, or require task-specific reward labels to facilitate policy optimization via Reinforcement Learning (RL). They are costly due to the substantial human efforts required to collect expert data or design reward functions. To address these challenges, we aim to develop a versatile diffusion planner capable of leveraging large-scale inferior data that contains taskagnostic sub-optimal trajectories, with the ability to fast adapt to specific tasks. In this paper, we propose SODP, a two-stage framework that leverages Sub-Optimal data to learn a Diffusion Planner, which is generalizable for various downstream tasks. Specifically, in the pre-training stage, we train a foundation diffusion planner that extracts general planning capabilities by modeling the versatile distribution of multi-task trajectories, which can be sub-optimal and has wide data coverage. Then for downstream tasks, we adopt RLbased fine-tuning with task-specific rewards to quickly refine the diffusion planner, which aims to generate action sequences with higher taskspecific returns. Experimental results from multitask domains including Meta-World and Adroit demonstrate that SODP outperforms state-of-theart methods with only a small amount of data for reward-guided fine-tuning.

Original languageEnglish
Pages (from-to)15677-15699
Number of pages23
JournalProceedings of Machine Learning Research
Volume267
StatePublished - 2025
Event42nd International Conference on Machine Learning, ICML 2025 - Vancouver, Canada
Duration: 13 Jul 202519 Jul 2025

Fingerprint

Dive into the research topics of 'Task-Agnostic Pre-training and Task-Guided Fine-tuning for Versatile Diffusion Planner'. Together they form a unique fingerprint.

Cite this