Abstract
Learning diverse goal-conditioned tasks under sparse rewards remains a major challenge in deep reinforcement learning (DRL), particularly for underwater robotic manipulation where dynamic disturbances hinder exploration and exacerbate hindsight bias. To address these challenges, this article proposes an integrated framework that combines a temporal-context attention-based adaptive intrinsic curiosity (TCA-AIC) mechanism with an importance-weighted similarity-based goal hindsight experience replay (IS-GHER) strategy. The TCA-AIC module enhances exploration efficiency and robustness to underwater disturbances by leveraging multistep temporal context and channel-wise attention in state prediction. Meanwhile, IS-GHER mitigates policy bias and stabilizes Q-value estimation by adaptively weighting hindsight relabeling samples based on goal similarity and training dynamics. The proposed framework is validated on several binaryreward underwater manipulation tasks in Gazebo simulation and further tested on a physical robot platform. Results show that it significantly improves learning efficiency and convergence speed compared with state-of-the-art baselines, particularly under sparse-reward and dynamically perturbed conditions.
| Original language | English |
|---|---|
| Journal | IEEE Transactions on Industrial Electronics |
| DOIs | |
| State | Accepted/In press - 2026 |
Keywords
- Adaptive intrinsic curiosity (AIC)
- deep reinforcement learning (DRL)
- hindsight experience replay (HER)
- underwater manipulation
Fingerprint
Dive into the research topics of 'Goal-Conditioned Reinforcement Learning-Based Control for Multihorizon Underwater Manipulation Under Sparse Rewards'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver