activity
20242026
collaborators

9 papers

cs.DC2026

LASER: Load-Aware Serving with Early-Exit for Reasoning LLMs at the Edge

Zhiqing Tang, Size Li, Hanshuai Cui +5

Large reasoning models (LRMs) such as DeepSeek-R1 have achieved strong performance through extended chain-of-thought (CoT) generation. However, deploying them on edge devices raise…

cs.AI2026

HeraSys: Collaborative Serving of Multiple LLM Workflows via Fine-Grained End-to-End Optimization

Size Li, Zhiqing Tang, Hongrui Liang +4

The proliferation of Large Language Models (LLMs) has shifted serving systems from processing isolated requests to orchestrating high-concurrency, multi-tenant agentic workflows. H…

cs.LG2025

Diffusion Models for Reinforcement Learning: Foundations, Taxonomy, and Development

Changfu Xu, Jianxiong Guo, Yuzhu Liang +7

Diffusion Models (DMs), as a leading class of generative models, offer key advantages for reinforcement learning (RL), including multi-modal expressiveness, stable training, and tr…

cs.AI2025

Adaptive AI Agent Placement and Migration in Edge Intelligence Systems

Xingdan Wang, Jiayi He, Zhiqing Tang +5

The rise of LLMs such as ChatGPT and Claude fuels the need for AI agents capable of real-time task handling. However, migrating data-intensive, multi-modal edge workloads to cloud…

cs.DC2025

LRScheduler: A Layer-aware and Resource-adaptive Container Scheduler in Edge Computing

Zhiqing Tang, Wentao Peng, Jianxiong Guo +5

Lightweight containers provide an efficient approach for deploying computation-intensive applications in network edge. The layered storage structure of container images can further…

cs.DC2025

Adaptive Configuration Selection for Multi-Model Inference Pipelines in Edge Computing

Jinhao Sheng, Zhiqing Tang, Jianxiong Guo +1

The growing demand for real-time processing tasks is driving the need for multi-model inference pipelines on edge devices. However, cost-effectively deploying these pipelines while…