activity
20242026
most citedCOSMIC: Enabling Full-Stack Co-Design and Optimization of Distributed Machine Learning Systems

1 citations · 1 across the 3 of their papers we have counts for

collaborators

5 papers

cs.DC2026

Flint: Compiler Enabled Cluster-Free Design Space Exploration for Distributed ML

Jinsun Yoo, Meghan Cowan, Zheng Du +3

Design space exploration for future distributed Machine Learning systems suffers from a lack of readily available workload representation that enables flexible exploration across t…

cs.DC2026

Evaluating Cross-Architecture Performance Modeling of Distributed ML Workloads Using StableHLO

Jonas Svedas, Nathan Laubeuf, Ryan Harvey +6

Predicting the performance of large-scale distributed machine learning (ML) workloads across multiple accelerator architectures remains a central challenge in ML system design. Exi…

cs.AR2026

SCALE-Sim TPU: Validating and Extending SCALE-Sim for TPUs

Jingtian Dang, Ritik Raj, Changhai Man +2

Cycle-accurate simulators are widely used to study systolic accelerators, yet their accuracy and usability are often limited by weak validation against real hardware and poor integ…

cs.DC20251 cited

COSMIC: Enabling Full-Stack Co-Design and Optimization of Distributed Machine Learning Systems

Aditi Raju, Jared Ni, William Won +6

Large-scale machine learning models necessitate distributed systems, posing significant design challenges due to the large parameter space across distinct design stacks. Existing s…

cs.LG2024

LayerDAG: A Layerwise Autoregressive Diffusion Model for Directed Acyclic Graph Generation

Mufei Li, Viraj Shitole, Eli Chien +6

Directed acyclic graphs (DAGs) serve as crucial data representations in domains such as hardware synthesis and compiler/program optimization for computing systems. DAG generative m…