跳到主要导航 跳到搜索 跳到主要内容

Controlling Value Function Bias: A Novel Approach for Continuous State-Action Space Reinforcement Learning

  • Northwestern Polytechnical University Xian

科研成果: 期刊稿件文章同行评审

摘要

Reinforcementlearning (RL) has demonstrated remarkable performance in complex decision-making tasks. However, value function approximation based on neural networks commonly suffers from overestimation bias. While existing research mitigates overestimation by introducing multiple critic networks, this strategy often leads to underestimation bias. To address the above issues, this article proposes two novel algorithms based on the deep deterministic policy gradient (DDPG) framework. First, to reduce value estimation variance and alleviate overestimation, we propose the average deep deterministic policy gradient (ADDPG) algorithm, which estimates the action-value function by averaging the outputs of different critic networks. Then, to address the underestimation phenomenon that may arise from using two critic networks, we propose the moving average deep deterministic policy gradient (MOADDPG) algorithm, whose objective function is dynamically updated during training to balance estimation bias and adapt to the learning process. We provide a theoretical convergence analysis of both algorithms, demonstrating their advantages in stability and convergence. Finally, we conduct experiments on the benchmark platform. Simulation results show that the final performance of all proposed algorithms outperforms the baseline methods.

源语言英语
期刊IEEE Transactions on Industrial Informatics
DOI
出版状态已接受/待刊 - 2026

学术指纹

探究 'Controlling Value Function Bias: A Novel Approach for Continuous State-Action Space Reinforcement Learning' 的科研主题。它们共同构成独一无二的学术指纹。

引用此