Skip to main navigation Skip to search Skip to main content

Resource-Aware Surprise Reinforcement Learning for Collision Avoidance in Maritime UAV Encounters

  • Northwestern Polytechnical University Xian
  • City University of Hong Kong

Research output: Contribution to journalArticlepeer-review

Abstract

Collision avoidance in maritime unmanned aerial vehicle (UAV) operations must satisfy two competing objectives: ensuring reliable safety separation and minimizing unnecessary maneuver commands that increase operator burden and communication overhead. While deep reinforcement learning (DRL) has shown promise in handling high-dimensional encounter states, standard DRL approaches often prioritize safety at the cost of operational suitability, leading to frequent, oscillatory, or unnecessary avoidance commands that erode remote operator trust and consume limited communication bandwidth. To address this challenge, this paper proposes Resource-Aware Intrinsic Surprise Exploration (RAISE), a unified framework that balances collision avoidance performance with command economy. We conceptualize the issuance of avoidance maneuvers as a consumable “virtual resource”, compelling the agent to optimize its intervention budget. RAISE integrates this mechanism into the Soft Actor–Critic (SAC) architecture, augmented by a surprise-based intrinsic reward derived from the ensemble forward dynamics prediction error. This allows the agent to efficiently explore complex encounter scenarios driven by curiosity, while a resource-aware coefficient adaptively suppresses redundant actions when the communication or operational budget is constrained. Furthermore, an adaptive exponential moving average (EMA) scaling mechanism is introduced to stabilize the interplay between intrinsic and extrinsic rewards. Extensive simulations under diverse resource constraints and encounter geometries demonstrate that RAISE outperforms state-of-the-art baselines. It significantly reduces maneuver reversal rates and strengthens command stability without compromising safety margins. Specifically, under resource-constrained settings, RAISE suppresses excessive and unstable advisory behavior by reducing strengthening and reversal commands while maintaining effective collision avoidance; under resource-rich settings, it flexibly enhances safety buffers, demonstrating superior adaptability and operational realism for autonomous maritime UAV systems. Robustness evaluation confirms that RAISE maintains stable performance under sensor noise and wind disturbances.

Original languageEnglish
Article number450
JournalDrones
Volume10
Issue number6
DOIs
StatePublished - Jun 2026

Keywords

  • autonomous UAV navigation
  • deep reinforcement learning
  • intrinsic motivation
  • maritime UAV collision avoidance
  • resource-aware exploration

Fingerprint

Dive into the research topics of 'Resource-Aware Surprise Reinforcement Learning for Collision Avoidance in Maritime UAV Encounters'. Together they form a unique fingerprint.

Cite this