5 papers · 1 filter
LASER: Load-Aware Serving with Early-Exit for Reasoning LLMs at the Edge
Zhiqing Tang, Size Li, Hanshuai Cui +5
Large reasoning models (LRMs) such as DeepSeek-R1 have achieved strong performance through extended chain-of-thought (CoT) generation. However, deploying them on edge devices raise…
RISE: Relay Inference and Online Scheduling for Efficient Edge-Device Collaborative Diffusion Model Services
Zilan Huang, Zhiqing Tang, Hanshuai Cui +4
Text-to-image diffusion models are increasingly deployed at the network edge to serve heterogeneous workloads with diverse quality and latency requirements. However, existing deplo…
EAT: QoS-Aware Edge-Collaborative AIGC Task Scheduling via Attention-Guided Diffusion Reinforcement Learning
Zhifei Xu, Zhiqing Tang, Jiong Lou +5
The growth of Artificial Intelligence (AI) and large language models has enabled the use of Generative AI (GenAI) in cloud data centers for diverse AI-Generated Content (AIGC) task…
LRScheduler: A Layer-aware and Resource-adaptive Container Scheduler in Edge Computing
Zhiqing Tang, Wentao Peng, Jianxiong Guo +5
Lightweight containers provide an efficient approach for deploying computation-intensive applications in network edge. The layered storage structure of container images can further…
Adaptive Configuration Selection for Multi-Model Inference Pipelines in Edge Computing
Jinhao Sheng, Zhiqing Tang, Jianxiong Guo +1
The growing demand for real-time processing tasks is driving the need for multi-model inference pipelines on edge devices. However, cost-effectively deploying these pipelines while…