摘要
In the reinforcement learning (RL) system, one important issue is the tradeoff problem between exploration and exploitation. In this paper, we studied this dilemma and proposed a new approach to solving this problem by multiple-attribute decision making (MADM). The applicability of the proposed method is extended by transfer learning. The method decomposes a task into several subtasks and uses the policies of subtasks trained by RL. The proposed visual MADM method (V-MADM) is based on the state-action values of each subtask to select the action with maximal one. Meanwhile, this paper proposes a transfer learning method using a decay function with decreasing probability such that the prior experiences of the subtasks can be utilized to accelerate the learning rate. Finally, the experiment of robot confrontation and Maze walker is performed to evaluate the learning performance of the proposed method. The experimental results show that fewer training cost is needed to obtain a more effective learning performance.
| 源语言 | 英语 |
|---|---|
| 文章编号 | 8745507 |
| 页(从-至) | 695-708 |
| 页数 | 14 |
| 期刊 | IEEE Transactions on Cognitive and Developmental Systems |
| 卷 | 12 |
| 期 | 4 |
| DOI | |
| 出版状态 | 已出版 - 12月 2020 |
学术指纹
探究 'A multiple-attribute decision-making approach to reinforcement learning' 的科研主题。它们共同构成独一无二的学术指纹。引用此
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver