摘要
In this paper, an off-policy hierarchical reinforcement learning (HRL) algorithm is proposed to solve the collision avoidance problem for a class of multi-agent systems. The collision avoidance refers to maintaining a predefined formation pattern and avoiding collisions with obstacles while driving each agent to the target state, which is formulated as a differential game. We leverage the idea of divide and conquer to artificially decompose the problem into three corresponding subtasks: target state attraction, neighbor agent repulsion, and static obstacle repulsion, to cope with the complex external environment. The off-policy HRL algorithm is designed based on the original policy iteration algorithm and implemented in real-time using only measured data to cope with the problem of completely unknown system information. Compared with the traditional least-square and gradient descent approach, critic and action neural networks of each subtask are simultaneously added to a broad learning system (BLS). It is worth noting that the pseudo-inverse operation of BLS allows us to achieve a faster and better approximate solution of the weight using global data online. The uniform ultimate bounded stability of the closed-loop system is proved based on the Lyapunov approach. Finally, a simulation example is given to demonstrate the effectiveness of the developed algorithm.
| 源语言 | 英语 |
|---|---|
| 文章编号 | 108274 |
| 期刊 | Journal of the Franklin Institute |
| 卷 | 363 |
| 期 | 2 |
| DOI | |
| 出版状态 | 已出版 - 15 1月 2026 |
指纹
探究 'Off-policy hierarchical reinforcement learning for collision avoidance via broad learning system' 的科研主题。它们共同构成独一无二的指纹。引用此
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver