TY - JOUR
T1 - Learning Multiagent Cooperation via Reciprocity-Based Actor-Critic Framework
AU - Long, Jia
AU - Yu, Dengxiu
AU - Chen, C. L.Philip
N1 - Publisher Copyright:
© 2014 IEEE.
PY - 2026
Y1 - 2026
N2 - In multiagent mixed-motive games, the dynamic adaptation among agent policies induces a nonstationary environment, posing challenges to stable learning and convergence toward Pareto-optimal equilibrium. To address this, this article proposes a reciprocity-based actor-critic (AC) framework by embedding a mutual help mechanism, which encourages agents to balance self-interest with the preferences of others. Our approach utilizes a centralized critic to estimate Q-values and infer expected policies of other agents, while decentralized actors concurrently update individual policies through selective altruistic adjustments. By integrating expected policy computation and redefining the advantage function, it promotes agent cooperation without sacrificing individual interests, thus mitigating local optimality. A policy improvement lower bound is established to ensure monotonic convergence with theoretical guarantees for Pareto improvement. Furthermore, we design a hybrid neural architecture that combines a shared backbone network for global knowledge transfer with personalized branches for individual policy optimization. This structure balances coordination and adaptation by synthesizing extrinsic environmental feedback with intrinsic agent motivation. Experimental results validate the effectiveness of our method in enhancing behavior coordination and learning efficiency in decentralized multiagent settings.
AB - In multiagent mixed-motive games, the dynamic adaptation among agent policies induces a nonstationary environment, posing challenges to stable learning and convergence toward Pareto-optimal equilibrium. To address this, this article proposes a reciprocity-based actor-critic (AC) framework by embedding a mutual help mechanism, which encourages agents to balance self-interest with the preferences of others. Our approach utilizes a centralized critic to estimate Q-values and infer expected policies of other agents, while decentralized actors concurrently update individual policies through selective altruistic adjustments. By integrating expected policy computation and redefining the advantage function, it promotes agent cooperation without sacrificing individual interests, thus mitigating local optimality. A policy improvement lower bound is established to ensure monotonic convergence with theoretical guarantees for Pareto improvement. Furthermore, we design a hybrid neural architecture that combines a shared backbone network for global knowledge transfer with personalized branches for individual policy optimization. This structure balances coordination and adaptation by synthesizing extrinsic environmental feedback with intrinsic agent motivation. Experimental results validate the effectiveness of our method in enhancing behavior coordination and learning efficiency in decentralized multiagent settings.
KW - Hybrid neural architecture
KW - policy Pareto improvement
KW - reciprocity-based actor-critic (AC) framework
KW - selective altruistic coordination
UR - https://www.scopus.com/pages/publications/105043454758
U2 - 10.1109/TCSS.2026.3705509
DO - 10.1109/TCSS.2026.3705509
M3 - 文章
AN - SCOPUS:105043454758
SN - 2329-924X
JO - IEEE Transactions on Computational Social Systems
JF - IEEE Transactions on Computational Social Systems
ER -