TY - JOUR
T1 - Hierarchical Learning in Distributed Online Markov Games via Partial Cooperation
AU - Long, Jia
AU - Yu, Dengxiu
AU - Wang, Zhen
N1 - Publisher Copyright:
© 2025 IEEE.
PY - 2026
Y1 - 2026
N2 - How to achieve emergent behaviors under incomplete information is a major challenge in multi-agent game learning. In this paper, we propose a generalized partial cooperation framework to realize Nash equilibrium (NE) selection and transition in distributed online Markov games, requiring only local interaction and information sharing. By introducing the graph self-attention mechanism into policy hierarchy, the behavior logic of the agent is decomposed to collaboratively learn the global optimal NE point from two distinct time scales. This leads to the corresponding bilevel optimization problem: the upper-level fine-tunes the reward structure to eliminate suboptimal equilibrium, and the lower-level learns optimal policy within the reformulated game. For the non-convex problem with non-unique NE, we develop a novel algorithm by concurrently integrating distributed online optimization and learning theory in the non-stationary environment. Specifically, lower-level utilizes Q-learning to acquire optimal policy without any prior knowledge, while upper-level inherits the environmental information explored by lower-level and uses a distributed Alternating Direction Method of Multipliers (ADMM) to adjust reward-sharing weight. In addition, we give a convergence proof of the alternating learning and optimization iteration. Finally, simulations on the multi-agent prisoner's dilemma and Uncrewed Aerial Vehicle (UAV) coverage control task are presented to demonstrate the effectiveness of proposed algorithm.
AB - How to achieve emergent behaviors under incomplete information is a major challenge in multi-agent game learning. In this paper, we propose a generalized partial cooperation framework to realize Nash equilibrium (NE) selection and transition in distributed online Markov games, requiring only local interaction and information sharing. By introducing the graph self-attention mechanism into policy hierarchy, the behavior logic of the agent is decomposed to collaboratively learn the global optimal NE point from two distinct time scales. This leads to the corresponding bilevel optimization problem: the upper-level fine-tunes the reward structure to eliminate suboptimal equilibrium, and the lower-level learns optimal policy within the reformulated game. For the non-convex problem with non-unique NE, we develop a novel algorithm by concurrently integrating distributed online optimization and learning theory in the non-stationary environment. Specifically, lower-level utilizes Q-learning to acquire optimal policy without any prior knowledge, while upper-level inherits the environmental information explored by lower-level and uses a distributed Alternating Direction Method of Multipliers (ADMM) to adjust reward-sharing weight. In addition, we give a convergence proof of the alternating learning and optimization iteration. Finally, simulations on the multi-agent prisoner's dilemma and Uncrewed Aerial Vehicle (UAV) coverage control task are presented to demonstrate the effectiveness of proposed algorithm.
KW - Partial cooperation
KW - bilevel optimization
KW - graph self-attention mechanism
KW - policy hierarchy
UR - https://www.scopus.com/pages/publications/105020294257
U2 - 10.1109/TSIPN.2025.3626133
DO - 10.1109/TSIPN.2025.3626133
M3 - 文章
AN - SCOPUS:105020294257
SN - 2373-776X
VL - 12
SP - 681
EP - 691
JO - IEEE Transactions on Signal and Information Processing over Networks
JF - IEEE Transactions on Signal and Information Processing over Networks
ER -