跳到主要导航 跳到搜索 跳到主要内容

Hierarchical Meta-Graph Reinforcement Learning for Collaborative GenAI Model Caching and Inference Scheduling

  • Liang Zhao
  • , Jing Wei
  • , Huan Zhou
  • , Zhen Chen
  • , Tong Wu
  • , Victor C.M. Leung
  • China Three Gorges University
  • Jinan University
  • Zhejiang Gongshang University
  • University of British Columbia

科研成果: 期刊稿件文章同行评审

摘要

Enabling collaborative generative AI (GenAI) inference at the network edge is challenging due to limited caching capacity, heterogeneous computing resources, and highly dynamic, latency-sensitive service demands. In this paper, we investigate the joint optimization of GenAI model caching, inference offloading, and resource allocation in a collaborative cloud-edge-end architecture. To address the strong coupling between long-term caching decisions and short-term scheduling dynamics, we propose a Hierarchical Meta-Graph Reinforcement Learning framework, termed HMGRL. Specifically, a heat-greedy model caching strategy is developed to capture time-varying model popularity and to reduce switching overhead on a slow timescale, while a graph-enhanced dueling deep reinforcement learning algorithm with prioritized experience replay enables topology-aware collaborative inference offloading and resource allocation on a fast timescale. Extensive simulations demonstrate that HMGRL consistently outperforms representative baselines in terms of system utility, cache and computing-resource utilization, convergence stability, and performance robustness. These results validate the effectiveness of the proposed hierarchical learning framework for practical GenAI applications at the network edge.

源语言英语
页(从-至)10216-10231
页数16
期刊IEEE Transactions on Cognitive Communications and Networking
12
DOI
出版状态已出版 - 2026

学术指纹

探究 'Hierarchical Meta-Graph Reinforcement Learning for Collaborative GenAI Model Caching and Inference Scheduling' 的科研主题。它们共同构成独一无二的学术指纹。

引用此