Skip to main navigation Skip to search Skip to main content

Controlling Value Function Bias: A Novel Approach for Continuous State-Action Space Reinforcement Learning

  • Northwestern Polytechnical University Xian

Research output: Contribution to journalArticlepeer-review

Abstract

Reinforcementlearning (RL) has demonstrated remarkable performance in complex decision-making tasks. However, value function approximation based on neural networks commonly suffers from overestimation bias. While existing research mitigates overestimation by introducing multiple critic networks, this strategy often leads to underestimation bias. To address the above issues, this article proposes two novel algorithms based on the deep deterministic policy gradient (DDPG) framework. First, to reduce value estimation variance and alleviate overestimation, we propose the average deep deterministic policy gradient (ADDPG) algorithm, which estimates the action-value function by averaging the outputs of different critic networks. Then, to address the underestimation phenomenon that may arise from using two critic networks, we propose the moving average deep deterministic policy gradient (MOADDPG) algorithm, whose objective function is dynamically updated during training to balance estimation bias and adapt to the learning process. We provide a theoretical convergence analysis of both algorithms, demonstrating their advantages in stability and convergence. Finally, we conduct experiments on the benchmark platform. Simulation results show that the final performance of all proposed algorithms outperforms the baseline methods.

Original languageEnglish
JournalIEEE Transactions on Industrial Informatics
DOIs
StateAccepted/In press - 2026

Keywords

  • Deep deterministic policy gradient (DDPG)
  • deep reinforcement learning
  • overestimation problem
  • underestimation problem

Fingerprint

Dive into the research topics of 'Controlling Value Function Bias: A Novel Approach for Continuous State-Action Space Reinforcement Learning'. Together they form a unique fingerprint.

Cite this