TY - JOUR
T1 - Research on Decision-Making Strategies for Multi-Agent UAVs in Island Missions Based on Rainbow Fusion MADDPG Algorithm
AU - Yang, Chaofan
AU - Zhang, Bo
AU - Zhang, Meng
AU - Wang, Qi
AU - Zhu, Peican
N1 - Publisher Copyright:
© 2025 by the authors.
PY - 2025/10
Y1 - 2025/10
N2 - Highlights: What are the main findings? This study presents an enhanced algorithm that integrates the Rainbow module to improve the Multi-Agent Deep Deterministic Policy Gradient (MADDPG) algorithm for multi-agent UAV cooperative and competitive scenarios. The proposed algorithm incorporates Prioritized Experience Replay (PER) and multi-step TD updating to optimize long-term reward perception and enhance learning efficiency. Behavioral cloning is also employed to accelerate convergence during initial training. What is the implication of the main finding? Experimental results on a UAV island capture simulation demonstrate that the enhanced algorithm outperforms the original MADDPG, showing a 40% increase in convergence speed and a doubled combat power preservation rate. The algorithm proves to be a robust and efficient solution for complex, dynamic, multi-agent game environments. To address the limitations of the Multi-Agent Deep Deterministic Policy Gradient (MADDPG) algorithm in autonomous control tasks including low convergence efficiency, poor training stability, inadequate adaptability of confrontation strategies, and challenges in handling sparse reward tasks—this paper proposes an enhanced algorithm by integrating the Rainbow module. The proposed algorithm improves long-term reward optimization through prioritized experience replay (PER) and multi-step TD updating mechanisms. Additionally, a dynamic reward allocation strategy is introduced to enhance the collaborative and adaptive decision-making capabilities of agents in complex adversarial scenarios. Furthermore, behavioral cloning is employed to accelerate convergence during the initial training phase. Extensive experiments are conducted on the MaCA simulation platform for 5 vs. 5 to 10 vs. 10 UAV island capture missions. The results demonstrate that the Rainbow-MADDPG outperforms the original MADDPG in several key metrics: (1) The average reward value improves across all confrontation scales, with notable enhancements in 6 vs. 6 and 7 vs. 7 tasks, achieving reward values of 14, representing 6.05-fold and 2.5-fold improvements over the baseline, respectively. (2) The convergence speed increases by 40%. (3) The combat effectiveness preservation rate doubles that of the baseline. Moreover, the algorithm achieves the highest average reward value in quasi-rectangular island scenarios, demonstrating its strong adaptability to large-scale dynamic game environments. This study provides an innovative technical solution to address the challenges of strategy stability and efficiency imbalance in multi-agent autonomous control tasks, with significant application potential in UAV defense, cluster cooperative tasks, and related fields.
AB - Highlights: What are the main findings? This study presents an enhanced algorithm that integrates the Rainbow module to improve the Multi-Agent Deep Deterministic Policy Gradient (MADDPG) algorithm for multi-agent UAV cooperative and competitive scenarios. The proposed algorithm incorporates Prioritized Experience Replay (PER) and multi-step TD updating to optimize long-term reward perception and enhance learning efficiency. Behavioral cloning is also employed to accelerate convergence during initial training. What is the implication of the main finding? Experimental results on a UAV island capture simulation demonstrate that the enhanced algorithm outperforms the original MADDPG, showing a 40% increase in convergence speed and a doubled combat power preservation rate. The algorithm proves to be a robust and efficient solution for complex, dynamic, multi-agent game environments. To address the limitations of the Multi-Agent Deep Deterministic Policy Gradient (MADDPG) algorithm in autonomous control tasks including low convergence efficiency, poor training stability, inadequate adaptability of confrontation strategies, and challenges in handling sparse reward tasks—this paper proposes an enhanced algorithm by integrating the Rainbow module. The proposed algorithm improves long-term reward optimization through prioritized experience replay (PER) and multi-step TD updating mechanisms. Additionally, a dynamic reward allocation strategy is introduced to enhance the collaborative and adaptive decision-making capabilities of agents in complex adversarial scenarios. Furthermore, behavioral cloning is employed to accelerate convergence during the initial training phase. Extensive experiments are conducted on the MaCA simulation platform for 5 vs. 5 to 10 vs. 10 UAV island capture missions. The results demonstrate that the Rainbow-MADDPG outperforms the original MADDPG in several key metrics: (1) The average reward value improves across all confrontation scales, with notable enhancements in 6 vs. 6 and 7 vs. 7 tasks, achieving reward values of 14, representing 6.05-fold and 2.5-fold improvements over the baseline, respectively. (2) The convergence speed increases by 40%. (3) The combat effectiveness preservation rate doubles that of the baseline. Moreover, the algorithm achieves the highest average reward value in quasi-rectangular island scenarios, demonstrating its strong adaptability to large-scale dynamic game environments. This study provides an innovative technical solution to address the challenges of strategy stability and efficiency imbalance in multi-agent autonomous control tasks, with significant application potential in UAV defense, cluster cooperative tasks, and related fields.
KW - MADDPG
KW - multi-agent
KW - multi-step TD update
KW - prioritized experience replay
KW - rainbow
KW - reinforcement learning
UR - https://www.scopus.com/pages/publications/105020020455
U2 - 10.3390/drones9100673
DO - 10.3390/drones9100673
M3 - 文章
AN - SCOPUS:105020020455
SN - 2504-446X
VL - 9
JO - Drones
JF - Drones
IS - 10
M1 - 673
ER -