1 citations · 2 across the 6 of their papers we have counts for
5 papers · 1 filter
MLCommons Chakra: Advancing Performance Benchmarking and Co-design using Standardized Execution Traces
Srinivas Sridharan, Theodor-Adrian Badea, Andy Balogh +26
The fast pace of artificial intelligence~(AI) innovation demands an agile methodology for observation, reproduction and optimization of distributed machine learning~(ML) workload b…
Flint: Compiler Enabled Cluster-Free Design Space Exploration for Distributed ML
Jinsun Yoo, Meghan Cowan, Zheng Du +3
Design space exploration for future distributed Machine Learning systems suffers from a lack of readily available workload representation that enables flexible exploration across t…
Scalable Synthesis of distributed LLM workloads through Symbolic Tensor Graphs
Changhai Man, Joongun Park, Hanjiang Wu +3
Optimizing the performance of large language models (LLMs) on large-scale AI training and inference systems requires a scalable and expressive mechanism to model distributed worklo…
COSMIC: Enabling Full-Stack Co-Design and Optimization of Distributed Machine Learning Systems
Aditi Raju, Jared Ni, William Won +6
Large-scale machine learning models necessitate distributed systems, posing significant design challenges due to the large parameter space across distinct design stacks. Existing s…
Towards a Standardized Representation for Deep Learning Collective Algorithms
Jinsun Yoo, William Won, Meghan Cowan +4
The explosion of machine learning model size has led to its execution on distributed clusters at a very large scale. Many works have tried to optimize the process of producing coll…