9 papers
LASER: Load-Aware Serving with Early-Exit for Reasoning LLMs at the Edge
Zhiqing Tang, Size Li, Hanshuai Cui +5
Large reasoning models (LRMs) such as DeepSeek-R1 have achieved strong performance through extended chain-of-thought (CoT) generation. However, deploying them on edge devices raise…
HeraSys: Collaborative Serving of Multiple LLM Workflows via Fine-Grained End-to-End Optimization
Size Li, Zhiqing Tang, Hongrui Liang +4
The proliferation of Large Language Models (LLMs) has shifted serving systems from processing isolated requests to orchestrating high-concurrency, multi-tenant agentic workflows. H…
Diffusion Models for Reinforcement Learning: Foundations, Taxonomy, and Development
Changfu Xu, Jianxiong Guo, Yuzhu Liang +7
Diffusion Models (DMs), as a leading class of generative models, offer key advantages for reinforcement learning (RL), including multi-modal expressiveness, stable training, and tr…
Adaptive AI Agent Placement and Migration in Edge Intelligence Systems
Xingdan Wang, Jiayi He, Zhiqing Tang +5
The rise of LLMs such as ChatGPT and Claude fuels the need for AI agents capable of real-time task handling. However, migrating data-intensive, multi-modal edge workloads to cloud…
LRScheduler: A Layer-aware and Resource-adaptive Container Scheduler in Edge Computing
Zhiqing Tang, Wentao Peng, Jianxiong Guo +5
Lightweight containers provide an efficient approach for deploying computation-intensive applications in network edge. The layered storage structure of container images can further…
Adaptive Configuration Selection for Multi-Model Inference Pipelines in Edge Computing
Jinhao Sheng, Zhiqing Tang, Jianxiong Guo +1
The growing demand for real-time processing tasks is driving the need for multi-model inference pipelines on edge devices. However, cost-effectively deploying these pipelines while…