activity
20242026
most citedTowards High-Goodput LLM Serving with Prefill-decode Multiplexing

1 citations · 1 across the 12 of their papers we have counts for

collaborators
Showing cs.DCShow all

5 papers · 1 filter

cs.DC2026

GF-DiT: Scheduling Parallelism for Diffusion Transformer Serving

Xinwei Qiang, Yifan Hu, Shixuan Sun +6

Diffusion Transformers (DiTs) have become the dominant architecture for image and video generation, creating growing demand for efficient DiT serving. Existing systems assign each…

cs.DC2025

Harli: SLO-Aware Co-location of LLM Inference and PEFT-based Finetuning on Model-as-a-Service Platforms

Ao Xu, Han Zhao, Weihao Cui +7

Large language models (LLMs) are increasingly deployed under the Model-as-a-Service (MaaS) paradigm. To meet stringent quality-of-service (QoS) requirements, existing LLM serving s…

cs.DC2025

Towards Resource-Efficient Serverless LLM Inference with SLINFER

Chuhao Xu, Zijun Li, Quan Chen +3

The rise of LLMs has driven demand for private serverless deployments, characterized by moderate-sized models and infrequent requests. While existing serverless solutions follow ex…

cs.DC2024

Towards Fast Setup and High Throughput of GPU Serverless Computing

Han Zhao, Weihao Cui, Quan Chen +6

Integrating GPUs into serverless computing platforms is crucial for improving efficiency. However, existing solutions for GPU-enabled serverless computing platforms face two signif…

cs.DC2024

Arena: Efficiently Training Large Models via Dynamic Scheduling and Adaptive Parallelism Co-Design

Chunyu Xue, Weihao Cui, Quan Chen +10

Efficiently training large-scale models (LMs) in GPU clusters involves two separate avenues: inter-job dynamic scheduling and intra-job adaptive parallelism (AP). However, existing…