TY - JOUR
T1 - A Self-Evolving Multi-Agent Reinforcement Learning for Multi-Target Encirclement with UGV-UAV Cluster in Urban Scenarios
AU - Zhang, Ying
AU - Cai, Wangze
AU - Chen, Jinchao
AU - Liang, Xianguang
AU - Du, Xiaoyan
AU - Du, Chenglie
N1 - Publisher Copyright:
© 2014 IEEE.
PY - 2026
Y1 - 2026
N2 - Unmanned ground vehicles–unmanned aerial vehicles (UGVs–UAVs) cluster plays a pivotal role in target encirclement missions in urban scenarios. However, the target encirclement performance of this heterogeneous cluster highly depends on its collaboration performance. In this paper, a self-evolving multi-agent reinforcement learning (SEMARL) approach is proposed for the UGV–UAV cluster to address the multi-target encirclement task in urban scenarios. To accurately depict the multi-target encirclement task, the models of the UGV–UAV cluster and the encirclement constraints are first established. Then an actor-critic-based learning framework is constructed by integrating a primitive decomposition-based heterogene-ous-agent proximal policy optimization (PD-HAPPO) algorithm and a gate attention-based reward self-evolving (GARSE) strategy. The PD-HAPPO algorithm decomposes the policy to allow the heterogeneous agents to share a common actor network while maintaining their own critic networks. In addition, the policy loss function combines policy stability, value loss, and entropy regularization primitives to ensure the stable update of the policy, improve value estimation accuracy, and avoid falling into local optima. To enable SEMARL with better evolvability, the GARSE strategy is proposed to update the weight vector of the reward. The proposed SEMARL approach is validated and compared with three state-of-the-art (SOTA) methods in several urban scenarios. Compared with SOTA methods, the proposed method improves the average reward and encirclement success rate by more than 11% and 3.2%, respectively, while reducing the average number of encirclement steps by more than 15%.
AB - Unmanned ground vehicles–unmanned aerial vehicles (UGVs–UAVs) cluster plays a pivotal role in target encirclement missions in urban scenarios. However, the target encirclement performance of this heterogeneous cluster highly depends on its collaboration performance. In this paper, a self-evolving multi-agent reinforcement learning (SEMARL) approach is proposed for the UGV–UAV cluster to address the multi-target encirclement task in urban scenarios. To accurately depict the multi-target encirclement task, the models of the UGV–UAV cluster and the encirclement constraints are first established. Then an actor-critic-based learning framework is constructed by integrating a primitive decomposition-based heterogene-ous-agent proximal policy optimization (PD-HAPPO) algorithm and a gate attention-based reward self-evolving (GARSE) strategy. The PD-HAPPO algorithm decomposes the policy to allow the heterogeneous agents to share a common actor network while maintaining their own critic networks. In addition, the policy loss function combines policy stability, value loss, and entropy regularization primitives to ensure the stable update of the policy, improve value estimation accuracy, and avoid falling into local optima. To enable SEMARL with better evolvability, the GARSE strategy is proposed to update the weight vector of the reward. The proposed SEMARL approach is validated and compared with three state-of-the-art (SOTA) methods in several urban scenarios. Compared with SOTA methods, the proposed method improves the average reward and encirclement success rate by more than 11% and 3.2%, respectively, while reducing the average number of encirclement steps by more than 15%.
KW - multi-agent reinforcement learning
KW - Multi-target encirclement
KW - self-evolution
KW - unmanned aerial vehicles
KW - unmanned ground vehicles
KW - urban scenario
UR - https://www.scopus.com/pages/publications/105046327565
U2 - 10.1109/JIOT.2026.3718354
DO - 10.1109/JIOT.2026.3718354
M3 - 文章
AN - SCOPUS:105046327565
SN - 2327-4662
JO - IEEE Internet of Things Journal
JF - IEEE Internet of Things Journal
ER -