Showing cs.DCShow all
2 papers · 1 filter
cs.DC2026
ServeGen: Workload Characterization and Generation of Large Language Model Serving in Production
Yuxing Xiang, Xue Li, Kun Qian +3
With the widespread adoption of Large Language Models (LLMs), serving LLM inference requests has become an increasingly important task, attracting active research advancements. Pra…
cs.DC2026
Amoeba: Runtime Tensor Parallel Transformation for LLM Inference Services
Haoyu Chen, Xue Li, Kun Qian +3
In Large Language Model (LLM) inference services, it is challenging to make a parallelism strategy configuration, to efficiently process the requests of variance context lengths. R…