Abstract
Unmanned ground vehicles–unmanned aerial vehicles (UGVs–UAVs) cluster plays a pivotal role in target encirclement missions in urban scenarios. However, the target encirclement performance of this heterogeneous cluster highly depends on its collaboration performance. In this paper, a self-evolving multi-agent reinforcement learning (SEMARL) approach is proposed for the UGV–UAV cluster to address the multi-target encirclement task in urban scenarios. To accurately depict the multi-target encirclement task, the models of the UGV–UAV cluster and the encirclement constraints are first established. Then an actor-critic-based learning framework is constructed by integrating a primitive decomposition-based heterogene-ous-agent proximal policy optimization (PD-HAPPO) algorithm and a gate attention-based reward self-evolving (GARSE) strategy. The PD-HAPPO algorithm decomposes the policy to allow the heterogeneous agents to share a common actor network while maintaining their own critic networks. In addition, the policy loss function combines policy stability, value loss, and entropy regularization primitives to ensure the stable update of the policy, improve value estimation accuracy, and avoid falling into local optima. To enable SEMARL with better evolvability, the GARSE strategy is proposed to update the weight vector of the reward. The proposed SEMARL approach is validated and compared with three state-of-the-art (SOTA) methods in several urban scenarios. Compared with SOTA methods, the proposed method improves the average reward and encirclement success rate by more than 11% and 3.2%, respectively, while reducing the average number of encirclement steps by more than 15%.
| Original language | English |
|---|---|
| Journal | IEEE Internet of Things Journal |
| DOIs | |
| State | Accepted/In press - 2026 |
Keywords
- multi-agent reinforcement learning
- Multi-target encirclement
- self-evolution
- unmanned aerial vehicles
- unmanned ground vehicles
- urban scenario
Fingerprint
Dive into the research topics of 'A Self-Evolving Multi-Agent Reinforcement Learning for Multi-Target Encirclement with UGV-UAV Cluster in Urban Scenarios'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver