Skip to main navigation Skip to search Skip to main content

Learning Multiagent Cooperation via Reciprocity-Based Actor-Critic Framework

  • Northwestern Polytechnical University Xian
  • South China University of Technology

Research output: Contribution to journalArticlepeer-review

Abstract

In multiagent mixed-motive games, the dynamic adaptation among agent policies induces a nonstationary environment, posing challenges to stable learning and convergence toward Pareto-optimal equilibrium. To address this, this article proposes a reciprocity-based actor-critic (AC) framework by embedding a mutual help mechanism, which encourages agents to balance self-interest with the preferences of others. Our approach utilizes a centralized critic to estimate Q-values and infer expected policies of other agents, while decentralized actors concurrently update individual policies through selective altruistic adjustments. By integrating expected policy computation and redefining the advantage function, it promotes agent cooperation without sacrificing individual interests, thus mitigating local optimality. A policy improvement lower bound is established to ensure monotonic convergence with theoretical guarantees for Pareto improvement. Furthermore, we design a hybrid neural architecture that combines a shared backbone network for global knowledge transfer with personalized branches for individual policy optimization. This structure balances coordination and adaptation by synthesizing extrinsic environmental feedback with intrinsic agent motivation. Experimental results validate the effectiveness of our method in enhancing behavior coordination and learning efficiency in decentralized multiagent settings.

Original languageEnglish
JournalIEEE Transactions on Computational Social Systems
DOIs
StateAccepted/In press - 2026

Keywords

  • Hybrid neural architecture
  • policy Pareto improvement
  • reciprocity-based actor-critic (AC) framework
  • selective altruistic coordination

Fingerprint

Dive into the research topics of 'Learning Multiagent Cooperation via Reciprocity-Based Actor-Critic Framework'. Together they form a unique fingerprint.

Cite this