activity
20242026
most citedHOBBIT: A Mixed Precision Expert Offloading System for Fast MoE Inference

2 citations · 2 across the 6 of their papers we have counts for

collaborators

9 papers

cs.DC2026

PCR: A Prefetch-Enhanced Cache Reuse System for Low-Latency RAG Serving

Wenfeng Wang, Xiaofeng Hou, Peng Tang +5

Retrieval-Augmented Generation (RAG) systems enhance the performance of large language models (LLMs) by incorporating supplementary retrieved documents, enabling more accurate and…

cs.DC2026

MuxTune: Efficient Multi-Task LLM Fine-Tuning in Multi-Tenant Datacenters via Spatial-Temporal Backbone Multiplexing

Chunyu Xue, Yi Pan, Weihao Cui +4

Parameter-Efficient Fine-Tuning (PEFT) is widely applied as the backend of fine-tuning APIs for large language model (LLM) customization in datacenters. Service providers deploy se…

cs.LG2025

MoE-SpeQ: Speculative Quantized Decoding with Proactive Expert Prefetching and Offloading for Mixture-of-Experts

Wenfeng Wang, Jiacheng Liu, Xiaofeng Hou +5

The immense memory requirements of state-of-the-art Mixture-of-Experts (MoE) models present a significant challenge for inference, often exceeding the capacity of a single accelera…

cs.AI2025

Boosting Embodied AI Agents through Perception-Generation Disaggregation and Asynchronous Pipeline Execution

Shulai Zhang, Ao Xu, Quan Chen +6

Embodied AI systems operate in dynamic environments, requiring seamless integration of perception and generation modules to process high-frequency input and output demands. Traditi…

cs.DC2025

Kairos: Low-latency Multi-Agent Serving with Shared LLMs and Excessive Loads in the Public Cloud

Jinyuan Chen, Jiuchen Shi, Quan Chen +1

Multi-agent applications utilize the advanced capabilities of large language models (LLMs) for intricate task completion through agent collaboration in a workflow. Under this situa…

cs.DC2025

Towards Resource-Efficient Serverless LLM Inference with SLINFER

Chuhao Xu, Zijun Li, Quan Chen +3

The rise of LLMs has driven demand for private serverless deployments, characterized by moderate-sized models and infrequent requests. While existing serverless solutions follow ex…