Skip to main navigation Skip to search Skip to main content

Hierarchical Meta-Graph Reinforcement Learning for Collaborative GenAI Model Caching and Inference Scheduling

  • Liang Zhao
  • , Jing Wei
  • , Huan Zhou
  • , Zhen Chen
  • , Tong Wu
  • , Victor C.M. Leung
  • China Three Gorges University
  • Jinan University
  • Zhejiang Gongshang University
  • University of British Columbia

Research output: Contribution to journalArticlepeer-review

Abstract

Enabling collaborative generative AI (GenAI) inference at the network edge is challenging due to limited caching capacity, heterogeneous computing resources, and highly dynamic, latency-sensitive service demands. In this paper, we investigate the joint optimization of GenAI model caching, inference offloading, and resource allocation in a collaborative cloud-edge-end architecture. To address the strong coupling between long-term caching decisions and short-term scheduling dynamics, we propose a Hierarchical Meta-Graph Reinforcement Learning framework, termed HMGRL. Specifically, a heat-greedy model caching strategy is developed to capture time-varying model popularity and to reduce switching overhead on a slow timescale, while a graph-enhanced dueling deep reinforcement learning algorithm with prioritized experience replay enables topology-aware collaborative inference offloading and resource allocation on a fast timescale. Extensive simulations demonstrate that HMGRL consistently outperforms representative baselines in terms of system utility, cache and computing-resource utilization, convergence stability, and performance robustness. These results validate the effectiveness of the proposed hierarchical learning framework for practical GenAI applications at the network edge.

Original languageEnglish
Pages (from-to)10216-10231
Number of pages16
JournalIEEE Transactions on Cognitive Communications and Networking
Volume12
DOIs
StatePublished - 2026

Keywords

  • collaborative edge computing
  • Generative AI
  • graph attention networks
  • hierarchical reinforcement learning
  • inference offloading
  • model caching

Fingerprint

Dive into the research topics of 'Hierarchical Meta-Graph Reinforcement Learning for Collaborative GenAI Model Caching and Inference Scheduling'. Together they form a unique fingerprint.

Cite this