跳到主要导航 跳到搜索 跳到主要内容

Preference-Aware Task Routing for Edge-Cloud Hierarchical Large Language Model Inference

  • Northwestern Polytechnical University Xian
  • Harbin Engineering University
  • Zhejiang Gongshang University
  • Zhejiang University of Technology

科研成果: 书/报告/会议事项章节会议稿件同行评审

摘要

The hierarchical edge-cloud architecture is regarded as a promising solution for deploying large models at the edge, as it effectively alleviates resource and latency constraints on edge devices. However, it often employs static routing that overlooks the heterogeneity of task requests. Therefore, we propose Lyra, a preference-aware online routing framework to bridge this gap. Lyra synergizes edge SLMs and cloud LLMs, employing Lyapunov drift-plus-penalty theory to dynamically optimize the inference quality-latency trade-off without requiring future traffic knowledge. Experimental results demonstrate that Lyra can effectively adapt to dynamic workloads, thereby enhancing system throughput to meet diverse user demands.

源语言英语
主期刊名INFOCOM 2026 - IEEE Conference on Computer Communications
出版商Institute of Electrical and Electronics Engineers Inc.
ISBN(电子版)9798331549619
DOI
出版状态已出版 - 2026
活动2026 IEEE Conference on Computer Communications, INFOCOM 2026 - Tokyo, 日本
期限: 18 5月 202621 5月 2026

出版系列

姓名Proceedings - IEEE INFOCOM
ISSN(印刷版)0743-166X

会议

会议2026 IEEE Conference on Computer Communications, INFOCOM 2026
国家/地区日本
Tokyo
时期18/05/2621/05/26

学术指纹

探究 'Preference-Aware Task Routing for Edge-Cloud Hierarchical Large Language Model Inference' 的科研主题。它们共同构成独一无二的学术指纹。

引用此