Abstract
Collision avoidance in maritime unmanned aerial vehicle (UAV) operations must satisfy two competing objectives: ensuring reliable safety separation and minimizing unnecessary maneuver commands that increase operator burden and communication overhead. While deep reinforcement learning (DRL) has shown promise in handling high-dimensional encounter states, standard DRL approaches often prioritize safety at the cost of operational suitability, leading to frequent, oscillatory, or unnecessary avoidance commands that erode remote operator trust and consume limited communication bandwidth. To address this challenge, this paper proposes Resource-Aware Intrinsic Surprise Exploration (RAISE), a unified framework that balances collision avoidance performance with command economy. We conceptualize the issuance of avoidance maneuvers as a consumable “virtual resource”, compelling the agent to optimize its intervention budget. RAISE integrates this mechanism into the Soft Actor–Critic (SAC) architecture, augmented by a surprise-based intrinsic reward derived from the ensemble forward dynamics prediction error. This allows the agent to efficiently explore complex encounter scenarios driven by curiosity, while a resource-aware coefficient adaptively suppresses redundant actions when the communication or operational budget is constrained. Furthermore, an adaptive exponential moving average (EMA) scaling mechanism is introduced to stabilize the interplay between intrinsic and extrinsic rewards. Extensive simulations under diverse resource constraints and encounter geometries demonstrate that RAISE outperforms state-of-the-art baselines. It significantly reduces maneuver reversal rates and strengthens command stability without compromising safety margins. Specifically, under resource-constrained settings, RAISE suppresses excessive and unstable advisory behavior by reducing strengthening and reversal commands while maintaining effective collision avoidance; under resource-rich settings, it flexibly enhances safety buffers, demonstrating superior adaptability and operational realism for autonomous maritime UAV systems. Robustness evaluation confirms that RAISE maintains stable performance under sensor noise and wind disturbances.
| Original language | English |
|---|---|
| Article number | 450 |
| Journal | Drones |
| Volume | 10 |
| Issue number | 6 |
| DOIs | |
| State | Published - Jun 2026 |
Keywords
- autonomous UAV navigation
- deep reinforcement learning
- intrinsic motivation
- maritime UAV collision avoidance
- resource-aware exploration
Fingerprint
Dive into the research topics of 'Resource-Aware Surprise Reinforcement Learning for Collision Avoidance in Maritime UAV Encounters'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver