Skip to main navigation Skip to search Skip to main content

Preference-Aware Task Routing for Edge-Cloud Hierarchical Large Language Model Inference

  • Northwestern Polytechnical University Xian
  • Harbin Engineering University
  • Zhejiang Gongshang University
  • Zhejiang University of Technology

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

Abstract

The hierarchical edge-cloud architecture is regarded as a promising solution for deploying large models at the edge, as it effectively alleviates resource and latency constraints on edge devices. However, it often employs static routing that overlooks the heterogeneity of task requests. Therefore, we propose Lyra, a preference-aware online routing framework to bridge this gap. Lyra synergizes edge SLMs and cloud LLMs, employing Lyapunov drift-plus-penalty theory to dynamically optimize the inference quality-latency trade-off without requiring future traffic knowledge. Experimental results demonstrate that Lyra can effectively adapt to dynamic workloads, thereby enhancing system throughput to meet diverse user demands.

Original languageEnglish
Title of host publicationINFOCOM 2026 - IEEE Conference on Computer Communications
PublisherInstitute of Electrical and Electronics Engineers Inc.
ISBN (Electronic)9798331549619
DOIs
StatePublished - 2026
Event2026 IEEE Conference on Computer Communications, INFOCOM 2026 - Tokyo, Japan
Duration: 18 May 202621 May 2026

Publication series

NameProceedings - IEEE INFOCOM
ISSN (Print)0743-166X

Conference

Conference2026 IEEE Conference on Computer Communications, INFOCOM 2026
Country/TerritoryJapan
CityTokyo
Period18/05/2621/05/26

Keywords

  • Edge-Cloud Collaboration
  • Hierarchical Inference
  • Large Language Models
  • Lyapunov Optimization
  • Preference-Aware Routing

Fingerprint

Dive into the research topics of 'Preference-Aware Task Routing for Edge-Cloud Hierarchical Large Language Model Inference'. Together they form a unique fingerprint.

Cite this