跳到主要导航 跳到搜索 跳到主要内容

Hierarchical Learning in Distributed Online Markov Games via Partial Cooperation

  • Northwestern Polytechnical University Xian

科研成果: 期刊稿件文章同行评审

1 引用 (Scopus)

摘要

How to achieve emergent behaviors under incomplete information is a major challenge in multi-agent game learning. In this paper, we propose a generalized partial cooperation framework to realize Nash equilibrium (NE) selection and transition in distributed online Markov games, requiring only local interaction and information sharing. By introducing the graph self-attention mechanism into policy hierarchy, the behavior logic of the agent is decomposed to collaboratively learn the global optimal NE point from two distinct time scales. This leads to the corresponding bilevel optimization problem: the upper-level fine-tunes the reward structure to eliminate suboptimal equilibrium, and the lower-level learns optimal policy within the reformulated game. For the non-convex problem with non-unique NE, we develop a novel algorithm by concurrently integrating distributed online optimization and learning theory in the non-stationary environment. Specifically, lower-level utilizes Q-learning to acquire optimal policy without any prior knowledge, while upper-level inherits the environmental information explored by lower-level and uses a distributed Alternating Direction Method of Multipliers (ADMM) to adjust reward-sharing weight. In addition, we give a convergence proof of the alternating learning and optimization iteration. Finally, simulations on the multi-agent prisoner's dilemma and Uncrewed Aerial Vehicle (UAV) coverage control task are presented to demonstrate the effectiveness of proposed algorithm.

源语言英语
页(从-至)681-691
页数11
期刊IEEE Transactions on Signal and Information Processing over Networks
12
DOI
出版状态已出版 - 2026

学术指纹

探究 'Hierarchical Learning in Distributed Online Markov Games via Partial Cooperation' 的科研主题。它们共同构成独一无二的学术指纹。

引用此