TY - JOUR
T1 - Controlling Value Function Bias
T2 - A Novel Approach for Continuous State-Action Space Reinforcement Learning
AU - Qi, Chenyang
AU - Wang, Rizhong
AU - Li, Huiping
N1 - Publisher Copyright:
© 2005-2012 IEEE.
PY - 2026
Y1 - 2026
N2 - Reinforcementlearning (RL) has demonstrated remarkable performance in complex decision-making tasks. However, value function approximation based on neural networks commonly suffers from overestimation bias. While existing research mitigates overestimation by introducing multiple critic networks, this strategy often leads to underestimation bias. To address the above issues, this article proposes two novel algorithms based on the deep deterministic policy gradient (DDPG) framework. First, to reduce value estimation variance and alleviate overestimation, we propose the average deep deterministic policy gradient (ADDPG) algorithm, which estimates the action-value function by averaging the outputs of different critic networks. Then, to address the underestimation phenomenon that may arise from using two critic networks, we propose the moving average deep deterministic policy gradient (MOADDPG) algorithm, whose objective function is dynamically updated during training to balance estimation bias and adapt to the learning process. We provide a theoretical convergence analysis of both algorithms, demonstrating their advantages in stability and convergence. Finally, we conduct experiments on the benchmark platform. Simulation results show that the final performance of all proposed algorithms outperforms the baseline methods.
AB - Reinforcementlearning (RL) has demonstrated remarkable performance in complex decision-making tasks. However, value function approximation based on neural networks commonly suffers from overestimation bias. While existing research mitigates overestimation by introducing multiple critic networks, this strategy often leads to underestimation bias. To address the above issues, this article proposes two novel algorithms based on the deep deterministic policy gradient (DDPG) framework. First, to reduce value estimation variance and alleviate overestimation, we propose the average deep deterministic policy gradient (ADDPG) algorithm, which estimates the action-value function by averaging the outputs of different critic networks. Then, to address the underestimation phenomenon that may arise from using two critic networks, we propose the moving average deep deterministic policy gradient (MOADDPG) algorithm, whose objective function is dynamically updated during training to balance estimation bias and adapt to the learning process. We provide a theoretical convergence analysis of both algorithms, demonstrating their advantages in stability and convergence. Finally, we conduct experiments on the benchmark platform. Simulation results show that the final performance of all proposed algorithms outperforms the baseline methods.
KW - Deep deterministic policy gradient (DDPG)
KW - deep reinforcement learning
KW - overestimation problem
KW - underestimation problem
UR - https://www.scopus.com/pages/publications/105045786311
U2 - 10.1109/TII.2026.3710460
DO - 10.1109/TII.2026.3710460
M3 - 文章
AN - SCOPUS:105045786311
SN - 1551-3203
JO - IEEE Transactions on Industrial Informatics
JF - IEEE Transactions on Industrial Informatics
ER -