IOB: integrating optimization transfer and behavior transfer for multi-policy reuse

Siyuan Li; Hao Li; Jin Zhang; Zhen Wang; Peng Liu; Chongjie Zhang

doi:10.1007/s10458-023-09630-9

IOB: integrating optimization transfer and behavior transfer for multi-policy reuse

Siyuan Li, Hao Li, Jin Zhang, Zhen Wang, Peng Liu, Chongjie Zhang

School of Cybersecurity

Research output: Contribution to journal › Article › peer-review

2 Scopus citations

Abstract

Humans have the ability to reuse previously learned policies to solve new tasks quickly, and reinforcement learning (RL) agents can do the same by transferring knowledge from source policies to a related target task. Transfer RL methods can reshape the policy optimization objective (optimization transfer) or influence the behavior policy (behavior transfer) using source policies. However, selecting the appropriate source policy with limited samples to guide target policy learning has been a challenge. Previous methods introduce additional components, such as hierarchical policies or estimations of source policies’ value functions, which can lead to non-stationary policy optimization or heavy sampling costs, diminishing transfer effectiveness. To address this challenge, we propose a novel transfer RL method that selects the source policy without training extra components. Our method utilizes the Q function in the actor-critic framework to guide policy selection, choosing the source policy with the largest one-step improvement over the current target policy. We integrate optimization transfer and behavior transfer (IOB) by regularizing the learned policy to mimic the guidance policy and combining them as the behavior policy. This integration significantly enhances transfer effectiveness, surpasses state-of-the-art transfer RL baselines in benchmark tasks, and improves final performance and knowledge transferability in continual learning scenarios. Additionally, we show that our optimization transfer technique is guaranteed to improve target policy learning.

Original language	English
Article number	3
Journal	Autonomous Agents and Multi-Agent Systems
Volume	38
Issue number	1
DOIs	https://doi.org/10.1007/s10458-023-09630-9
State	Published - Jun 2024

Keywords

Behavior transfer
Multi-policy reuse
Optimization transfer
Reinforcement learning

Access to Document

10.1007/s10458-023-09630-9

Cite this

@article{208fc419deae43089ecf4e0e77fa839e,

title = "IOB: integrating optimization transfer and behavior transfer for multi-policy reuse",

abstract = "Humans have the ability to reuse previously learned policies to solve new tasks quickly, and reinforcement learning (RL) agents can do the same by transferring knowledge from source policies to a related target task. Transfer RL methods can reshape the policy optimization objective (optimization transfer) or influence the behavior policy (behavior transfer) using source policies. However, selecting the appropriate source policy with limited samples to guide target policy learning has been a challenge. Previous methods introduce additional components, such as hierarchical policies or estimations of source policies{\textquoteright} value functions, which can lead to non-stationary policy optimization or heavy sampling costs, diminishing transfer effectiveness. To address this challenge, we propose a novel transfer RL method that selects the source policy without training extra components. Our method utilizes the Q function in the actor-critic framework to guide policy selection, choosing the source policy with the largest one-step improvement over the current target policy. We integrate optimization transfer and behavior transfer (IOB) by regularizing the learned policy to mimic the guidance policy and combining them as the behavior policy. This integration significantly enhances transfer effectiveness, surpasses state-of-the-art transfer RL baselines in benchmark tasks, and improves final performance and knowledge transferability in continual learning scenarios. Additionally, we show that our optimization transfer technique is guaranteed to improve target policy learning.",

keywords = "Behavior transfer, Multi-policy reuse, Optimization transfer, Reinforcement learning",

author = "Siyuan Li and Hao Li and Jin Zhang and Zhen Wang and Peng Liu and Chongjie Zhang",

note = "Publisher Copyright: {\textcopyright} 2023, Springer Science+Business Media, LLC, part of Springer Nature.",

year = "2024",

month = jun,

doi = "10.1007/s10458-023-09630-9",

language = "英语",

volume = "38",

journal = "Autonomous Agents and Multi-Agent Systems",

issn = "1387-2532",

publisher = "Springer Netherlands",

number = "1",

}

TY - JOUR

T1 - IOB

T2 - integrating optimization transfer and behavior transfer for multi-policy reuse

AU - Li, Siyuan

AU - Li, Hao

AU - Zhang, Jin

AU - Wang, Zhen

AU - Liu, Peng

AU - Zhang, Chongjie

PY - 2024/6

Y1 - 2024/6

N2 - Humans have the ability to reuse previously learned policies to solve new tasks quickly, and reinforcement learning (RL) agents can do the same by transferring knowledge from source policies to a related target task. Transfer RL methods can reshape the policy optimization objective (optimization transfer) or influence the behavior policy (behavior transfer) using source policies. However, selecting the appropriate source policy with limited samples to guide target policy learning has been a challenge. Previous methods introduce additional components, such as hierarchical policies or estimations of source policies’ value functions, which can lead to non-stationary policy optimization or heavy sampling costs, diminishing transfer effectiveness. To address this challenge, we propose a novel transfer RL method that selects the source policy without training extra components. Our method utilizes the Q function in the actor-critic framework to guide policy selection, choosing the source policy with the largest one-step improvement over the current target policy. We integrate optimization transfer and behavior transfer (IOB) by regularizing the learned policy to mimic the guidance policy and combining them as the behavior policy. This integration significantly enhances transfer effectiveness, surpasses state-of-the-art transfer RL baselines in benchmark tasks, and improves final performance and knowledge transferability in continual learning scenarios. Additionally, we show that our optimization transfer technique is guaranteed to improve target policy learning.

AB - Humans have the ability to reuse previously learned policies to solve new tasks quickly, and reinforcement learning (RL) agents can do the same by transferring knowledge from source policies to a related target task. Transfer RL methods can reshape the policy optimization objective (optimization transfer) or influence the behavior policy (behavior transfer) using source policies. However, selecting the appropriate source policy with limited samples to guide target policy learning has been a challenge. Previous methods introduce additional components, such as hierarchical policies or estimations of source policies’ value functions, which can lead to non-stationary policy optimization or heavy sampling costs, diminishing transfer effectiveness. To address this challenge, we propose a novel transfer RL method that selects the source policy without training extra components. Our method utilizes the Q function in the actor-critic framework to guide policy selection, choosing the source policy with the largest one-step improvement over the current target policy. We integrate optimization transfer and behavior transfer (IOB) by regularizing the learned policy to mimic the guidance policy and combining them as the behavior policy. This integration significantly enhances transfer effectiveness, surpasses state-of-the-art transfer RL baselines in benchmark tasks, and improves final performance and knowledge transferability in continual learning scenarios. Additionally, we show that our optimization transfer technique is guaranteed to improve target policy learning.

KW - Behavior transfer

KW - Multi-policy reuse

KW - Optimization transfer

KW - Reinforcement learning

UR - http://www.scopus.com/inward/record.url?scp=85179131469&partnerID=8YFLogxK

U2 - 10.1007/s10458-023-09630-9

DO - 10.1007/s10458-023-09630-9

M3 - 文章

AN - SCOPUS:85179131469

SN - 1387-2532

VL - 38

JO - Autonomous Agents and Multi-Agent Systems

JF - Autonomous Agents and Multi-Agent Systems

IS - 1

M1 - 3

ER -

IOB: integrating optimization transfer and behavior transfer for multi-policy reuse

Abstract

Keywords

Access to Document

Other files and links

Fingerprint

Cite this