跳到主要导航 跳到搜索 跳到主要内容

A Self-Evolving Multi-Agent Reinforcement Learning for Multi-Target Encirclement with UGV-UAV Cluster in Urban Scenarios

  • Ying Zhang
  • , Wangze Cai
  • , Jinchao Chen
  • , Xianguang Liang
  • , Xiaoyan Du
  • , Chenglie Du
  • Northwestern Polytechnical University Xian
  • Shaanxi Key Laboratory of Intelligent Policing

科研成果: 期刊稿件文章同行评审

摘要

Unmanned ground vehicles–unmanned aerial vehicles (UGVs–UAVs) cluster plays a pivotal role in target encirclement missions in urban scenarios. However, the target encirclement performance of this heterogeneous cluster highly depends on its collaboration performance. In this paper, a self-evolving multi-agent reinforcement learning (SEMARL) approach is proposed for the UGV–UAV cluster to address the multi-target encirclement task in urban scenarios. To accurately depict the multi-target encirclement task, the models of the UGV–UAV cluster and the encirclement constraints are first established. Then an actor-critic-based learning framework is constructed by integrating a primitive decomposition-based heterogene-ous-agent proximal policy optimization (PD-HAPPO) algorithm and a gate attention-based reward self-evolving (GARSE) strategy. The PD-HAPPO algorithm decomposes the policy to allow the heterogeneous agents to share a common actor network while maintaining their own critic networks. In addition, the policy loss function combines policy stability, value loss, and entropy regularization primitives to ensure the stable update of the policy, improve value estimation accuracy, and avoid falling into local optima. To enable SEMARL with better evolvability, the GARSE strategy is proposed to update the weight vector of the reward. The proposed SEMARL approach is validated and compared with three state-of-the-art (SOTA) methods in several urban scenarios. Compared with SOTA methods, the proposed method improves the average reward and encirclement success rate by more than 11% and 3.2%, respectively, while reducing the average number of encirclement steps by more than 15%.

源语言英语
期刊IEEE Internet of Things Journal
DOI
出版状态已接受/待刊 - 2026

学术指纹

探究 'A Self-Evolving Multi-Agent Reinforcement Learning for Multi-Target Encirclement with UGV-UAV Cluster in Urban Scenarios' 的科研主题。它们共同构成独一无二的学术指纹。

引用此