2 papers
cs.DC2025
Staggered Batch Scheduling: Co-optimizing Time-to-First-Token and Throughput for High-Efficiency LLM Inference
Jian Tian, Shuailong Li, Yang Cao +8
The evolution of Large Language Model (LLM) serving towards complex, distributed architectures--specifically the P/D-separated, large-scale DP+EP paradigm--introduces distinct sche…
cs.DC2025
Astra: Efficient and Money-saving Automatic Parallel Strategies Search on Heterogeneous GPUs
Peiran Wang, Haibing Li, Fu Haohan +3
In this paper, we introduce an efficient and money-saving automatic parallel strategies search framework on heterogeneous GPUs: Astra. First, Astra searches for the efficiency-opti…