2 papers
cs.DC2025
MoFa: A Unified Performance Modeling Framework for LLM Pretraining
Lu Zhao, Rong Shi, Shaoqing Zhang +14
The exponential growth in LLM scales, with parameters soaring from billions to trillions, has necessitated distributed pretraining across large clusters comprising thousands to ten…
cs.DC2025
Kant: An Efficient Unified Scheduling System for Large-Scale AI Clusters
Lingling Zeng, Gen Zhang, Jialin Peng +3
As AI cluster sizes continue to expand and the demand for large-language-model (LLM) training and inference workloads grows rapidly, traditional scheduling systems face significant…