8 papers
Scalable Synthesis of distributed LLM workloads through Symbolic Tensor Graphs
Changhai Man, Joongun Park, Hanjiang Wu +3
Optimizing the performance of large language models (LLMs) on large-scale AI training and inference systems requires a scalable and expressive mechanism to model distributed worklo…
ASTRA-sim 3.0: Next-Level Distributed Machine Learning Simulations via High-Fidelity GPU and Infrastructure Modeling
William Won, Jinsun Yoo, Tuan Ta +16
Distributed machine learning (ML) is a key paradigm for today's large-scale artificial intelligence applications. As model inference arises as an important use case, faithful model…
MLCommons Chakra: Advancing Performance Benchmarking and Co-design using Standardized Execution Traces
Srinivas Sridharan, Theodor-Adrian Badea, Andy Balogh +26
The fast pace of artificial intelligence~(AI) innovation demands an agile methodology for observation, reproduction and optimization of distributed machine learning~(ML) workload b…
Flint: Compiler Enabled Cluster-Free Design Space Exploration for Distributed ML
Jinsun Yoo, Meghan Cowan, Zheng Du +3
Design space exploration for future distributed Machine Learning systems suffers from a lack of readily available workload representation that enables flexible exploration across t…
Evaluating Cross-Architecture Performance Modeling of Distributed ML Workloads Using StableHLO
Jonas Svedas, Nathan Laubeuf, Ryan Harvey +6
Predicting the performance of large-scale distributed machine learning (ML) workloads across multiple accelerator architectures remains a central challenge in ML system design. Exi…
SCALE-Sim TPU: Validating and Extending SCALE-Sim for TPUs
Jingtian Dang, Ritik Raj, Changhai Man +2
Cycle-accurate simulators are widely used to study systolic accelerators, yet their accuracy and usability are often limited by weak validation against real hardware and poor integ…