跳到主要导航 跳到搜索 跳到主要内容

On the Value of Myopic Behavior in Policy Reuse

  • China Telecommunications
  • Fudan University
  • Guangxi Normal University
  • City University of Hong Kong
  • Hong Kong University of Science and Technology

科研成果: 期刊稿件文章同行评审

摘要

Leveraging learned strategies in unfamiliar scenarios is fundamental to human intelligence. In reinforcement learning, rationally reusing the policies acquired from other tasks or human experts is critical for tackling problems that are difficult to learn from scratch. In this work, we present a framework called Selective Myopic bEhavior Control (SMEC), which results from the insight that the short-term behaviors of prior policies are sharable across tasks. By evaluating the behaviors of prior policies via a hybrid value function architecture, SMEC adaptively aggregates the sharable short-term behaviors of prior policies and the long-term behaviors of the task policy, leading to coordinated decisions. Empirical results on a collection of manipulation and locomotion tasks demonstrate that SMEC outperforms existing methods, and validate the ability of SMEC to leverage related prior policies.

源语言英语
页(从-至)6647-6659
页数13
期刊IEEE Transactions on Pattern Analysis and Machine Intelligence
47
8
DOI
出版状态已出版 - 2025

学术指纹

探究 'On the Value of Myopic Behavior in Policy Reuse' 的科研主题。它们共同构成独一无二的学术指纹。

引用此