跳到主要导航 跳到搜索 跳到主要内容

Off-policy hierarchical reinforcement learning for collision avoidance via broad learning system

  • Tao Zhang
  • , Peixuan Song
  • , Jia Long
  • , Qian Kang
  • , Peng Li
  • , Dengxiu Yu
  • Northwestern Polytechnical University Xian
  • North Automatic Control Technology Institute

科研成果: 期刊稿件文章同行评审

摘要

In this paper, an off-policy hierarchical reinforcement learning (HRL) algorithm is proposed to solve the collision avoidance problem for a class of multi-agent systems. The collision avoidance refers to maintaining a predefined formation pattern and avoiding collisions with obstacles while driving each agent to the target state, which is formulated as a differential game. We leverage the idea of divide and conquer to artificially decompose the problem into three corresponding subtasks: target state attraction, neighbor agent repulsion, and static obstacle repulsion, to cope with the complex external environment. The off-policy HRL algorithm is designed based on the original policy iteration algorithm and implemented in real-time using only measured data to cope with the problem of completely unknown system information. Compared with the traditional least-square and gradient descent approach, critic and action neural networks of each subtask are simultaneously added to a broad learning system (BLS). It is worth noting that the pseudo-inverse operation of BLS allows us to achieve a faster and better approximate solution of the weight using global data online. The uniform ultimate bounded stability of the closed-loop system is proved based on the Lyapunov approach. Finally, a simulation example is given to demonstrate the effectiveness of the developed algorithm.

源语言英语
文章编号108274
期刊Journal of the Franklin Institute
363
2
DOI
出版状态已出版 - 15 1月 2026

指纹

探究 'Off-policy hierarchical reinforcement learning for collision avoidance via broad learning system' 的科研主题。它们共同构成独一无二的指纹。

引用此