2 papers
cs.DC2026
MLCommons Chakra: Advancing Performance Benchmarking and Co-design using Standardized Execution Traces
Srinivas Sridharan, Theodor-Adrian Badea, Andy Balogh +26
The fast pace of artificial intelligence~(AI) innovation demands an agile methodology for observation, reproduction and optimization of distributed machine learning~(ML) workload b…
cs.DC2025
Lumos: Efficient Performance Modeling and Estimation for Large-scale LLM Training
Mingyu Liang, Hiwot Tadese Kassa, Wenyin Fu +3
Training LLMs in distributed environments presents significant challenges due to the complexity of model execution, deployment systems, and the vast space of configurable strategie…