TY - JOUR
T1 - Resource-Aware Surprise Reinforcement Learning for Collision Avoidance in Maritime UAV Encounters
AU - Liu, Zuocheng
AU - Feng, Qi
AU - Wang, Zidong
AU - Gao, Xiaoguang
N1 - Publisher Copyright:
© 2026 by the authors.
PY - 2026/6
Y1 - 2026/6
N2 - Collision avoidance in maritime unmanned aerial vehicle (UAV) operations must satisfy two competing objectives: ensuring reliable safety separation and minimizing unnecessary maneuver commands that increase operator burden and communication overhead. While deep reinforcement learning (DRL) has shown promise in handling high-dimensional encounter states, standard DRL approaches often prioritize safety at the cost of operational suitability, leading to frequent, oscillatory, or unnecessary avoidance commands that erode remote operator trust and consume limited communication bandwidth. To address this challenge, this paper proposes Resource-Aware Intrinsic Surprise Exploration (RAISE), a unified framework that balances collision avoidance performance with command economy. We conceptualize the issuance of avoidance maneuvers as a consumable “virtual resource”, compelling the agent to optimize its intervention budget. RAISE integrates this mechanism into the Soft Actor–Critic (SAC) architecture, augmented by a surprise-based intrinsic reward derived from the ensemble forward dynamics prediction error. This allows the agent to efficiently explore complex encounter scenarios driven by curiosity, while a resource-aware coefficient adaptively suppresses redundant actions when the communication or operational budget is constrained. Furthermore, an adaptive exponential moving average (EMA) scaling mechanism is introduced to stabilize the interplay between intrinsic and extrinsic rewards. Extensive simulations under diverse resource constraints and encounter geometries demonstrate that RAISE outperforms state-of-the-art baselines. It significantly reduces maneuver reversal rates and strengthens command stability without compromising safety margins. Specifically, under resource-constrained settings, RAISE suppresses excessive and unstable advisory behavior by reducing strengthening and reversal commands while maintaining effective collision avoidance; under resource-rich settings, it flexibly enhances safety buffers, demonstrating superior adaptability and operational realism for autonomous maritime UAV systems. Robustness evaluation confirms that RAISE maintains stable performance under sensor noise and wind disturbances.
AB - Collision avoidance in maritime unmanned aerial vehicle (UAV) operations must satisfy two competing objectives: ensuring reliable safety separation and minimizing unnecessary maneuver commands that increase operator burden and communication overhead. While deep reinforcement learning (DRL) has shown promise in handling high-dimensional encounter states, standard DRL approaches often prioritize safety at the cost of operational suitability, leading to frequent, oscillatory, or unnecessary avoidance commands that erode remote operator trust and consume limited communication bandwidth. To address this challenge, this paper proposes Resource-Aware Intrinsic Surprise Exploration (RAISE), a unified framework that balances collision avoidance performance with command economy. We conceptualize the issuance of avoidance maneuvers as a consumable “virtual resource”, compelling the agent to optimize its intervention budget. RAISE integrates this mechanism into the Soft Actor–Critic (SAC) architecture, augmented by a surprise-based intrinsic reward derived from the ensemble forward dynamics prediction error. This allows the agent to efficiently explore complex encounter scenarios driven by curiosity, while a resource-aware coefficient adaptively suppresses redundant actions when the communication or operational budget is constrained. Furthermore, an adaptive exponential moving average (EMA) scaling mechanism is introduced to stabilize the interplay between intrinsic and extrinsic rewards. Extensive simulations under diverse resource constraints and encounter geometries demonstrate that RAISE outperforms state-of-the-art baselines. It significantly reduces maneuver reversal rates and strengthens command stability without compromising safety margins. Specifically, under resource-constrained settings, RAISE suppresses excessive and unstable advisory behavior by reducing strengthening and reversal commands while maintaining effective collision avoidance; under resource-rich settings, it flexibly enhances safety buffers, demonstrating superior adaptability and operational realism for autonomous maritime UAV systems. Robustness evaluation confirms that RAISE maintains stable performance under sensor noise and wind disturbances.
KW - autonomous UAV navigation
KW - deep reinforcement learning
KW - intrinsic motivation
KW - maritime UAV collision avoidance
KW - resource-aware exploration
UR - https://www.scopus.com/pages/publications/105042799466
U2 - 10.3390/drones10060450
DO - 10.3390/drones10060450
M3 - 文章
AN - SCOPUS:105042799466
SN - 2504-446X
VL - 10
JO - Drones
JF - Drones
IS - 6
M1 - 450
ER -