TY - JOUR
T1 - QFree
T2 - a universal value function factorisation for multi-agent reinforcement learning
AU - Wang, Rizhong
AU - Li, Huiping
AU - Cui, Di
AU - Xu, Demin
N1 - Publisher Copyright:
© 2026 Northeastern University, China.
PY - 2026
Y1 - 2026
N2 - Centralized training is widely utilised in the field of multi-agent reinforcement learning (MARL) to assure the stability of training process. Once a joint policy is obtained, it is critical to design a value function factorisation method to extract optimal decentralised policies for the agents, which needs to satisfy the individual-global-max (IGM) principle. While imposing additional limitations on the IGM function class can help to meet the requirement, it comes at the cost of restricting its application to more complex multi-agent environments. In this paper, we propose QFree, a universal value function factorisation method for MARL. We start by developing mathematical equivalent conditions of the IGM principle based on the advantage function, which ensures that the principle holds without any compromise, removing the conservatism of conventional methods. We then establish a more expressive mixing network architecture that can fulfill the equivalent factorisation. In particular, the novel loss function is developed by considering the equivalent conditions as regularisation term during policy evaluation in the MARL algorithm. Finally, the effectiveness of the proposed method is verified in a nonmonotonic matrix game scenario. Moreover, we show that QFree achieves the state-of-the-art performance in a general-purpose complex MARL benchmark environment, Starcraft Multi-Agent Challenge (SMAC).
AB - Centralized training is widely utilised in the field of multi-agent reinforcement learning (MARL) to assure the stability of training process. Once a joint policy is obtained, it is critical to design a value function factorisation method to extract optimal decentralised policies for the agents, which needs to satisfy the individual-global-max (IGM) principle. While imposing additional limitations on the IGM function class can help to meet the requirement, it comes at the cost of restricting its application to more complex multi-agent environments. In this paper, we propose QFree, a universal value function factorisation method for MARL. We start by developing mathematical equivalent conditions of the IGM principle based on the advantage function, which ensures that the principle holds without any compromise, removing the conservatism of conventional methods. We then establish a more expressive mixing network architecture that can fulfill the equivalent factorisation. In particular, the novel loss function is developed by considering the equivalent conditions as regularisation term during policy evaluation in the MARL algorithm. Finally, the effectiveness of the proposed method is verified in a nonmonotonic matrix game scenario. Moreover, we show that QFree achieves the state-of-the-art performance in a general-purpose complex MARL benchmark environment, Starcraft Multi-Agent Challenge (SMAC).
KW - Multi-agent reinforcement learning
KW - advantage function
KW - individual-global-max principle
KW - value function factorisation method
UR - https://www.scopus.com/pages/publications/105040384235
U2 - 10.1080/23307706.2025.2601691
DO - 10.1080/23307706.2025.2601691
M3 - 文章
AN - SCOPUS:105040384235
SN - 2330-7706
JO - Journal of Control and Decision
JF - Journal of Control and Decision
ER -