TY - JOUR
T1 - Goal-Conditioned Reinforcement Learning-Based Control for Multihorizon Underwater Manipulation Under Sparse Rewards
AU - Li, Yufeng
AU - Gao, Jian
AU - Chen, Yimin
AU - Liu, Jie
N1 - Publisher Copyright:
© 1982-2012 IEEE.
PY - 2026
Y1 - 2026
N2 - Learning diverse goal-conditioned tasks under sparse rewards remains a major challenge in deep reinforcement learning (DRL), particularly for underwater robotic manipulation where dynamic disturbances hinder exploration and exacerbate hindsight bias. To address these challenges, this article proposes an integrated framework that combines a temporal-context attention-based adaptive intrinsic curiosity (TCA-AIC) mechanism with an importance-weighted similarity-based goal hindsight experience replay (IS-GHER) strategy. The TCA-AIC module enhances exploration efficiency and robustness to underwater disturbances by leveraging multistep temporal context and channel-wise attention in state prediction. Meanwhile, IS-GHER mitigates policy bias and stabilizes Q-value estimation by adaptively weighting hindsight relabeling samples based on goal similarity and training dynamics. The proposed framework is validated on several binaryreward underwater manipulation tasks in Gazebo simulation and further tested on a physical robot platform. Results show that it significantly improves learning efficiency and convergence speed compared with state-of-the-art baselines, particularly under sparse-reward and dynamically perturbed conditions.
AB - Learning diverse goal-conditioned tasks under sparse rewards remains a major challenge in deep reinforcement learning (DRL), particularly for underwater robotic manipulation where dynamic disturbances hinder exploration and exacerbate hindsight bias. To address these challenges, this article proposes an integrated framework that combines a temporal-context attention-based adaptive intrinsic curiosity (TCA-AIC) mechanism with an importance-weighted similarity-based goal hindsight experience replay (IS-GHER) strategy. The TCA-AIC module enhances exploration efficiency and robustness to underwater disturbances by leveraging multistep temporal context and channel-wise attention in state prediction. Meanwhile, IS-GHER mitigates policy bias and stabilizes Q-value estimation by adaptively weighting hindsight relabeling samples based on goal similarity and training dynamics. The proposed framework is validated on several binaryreward underwater manipulation tasks in Gazebo simulation and further tested on a physical robot platform. Results show that it significantly improves learning efficiency and convergence speed compared with state-of-the-art baselines, particularly under sparse-reward and dynamically perturbed conditions.
KW - Adaptive intrinsic curiosity (AIC)
KW - deep reinforcement learning (DRL)
KW - hindsight experience replay (HER)
KW - underwater manipulation
UR - https://www.scopus.com/pages/publications/105045306506
U2 - 10.1109/TIE.2026.3708937
DO - 10.1109/TIE.2026.3708937
M3 - 文章
AN - SCOPUS:105045306506
SN - 0278-0046
JO - IEEE Transactions on Industrial Electronics
JF - IEEE Transactions on Industrial Electronics
ER -