activity
20162026
most citedMLPerf Training Benchmark

171 citations · 295 across the 54 of their papers we have counts for

collaborators
Showing cs.DCShow all

5 papers · 1 filter

cs.DC2026

MLSYSIM: First-Principles Infrastructure Modeling for Machine Learning Systems

Vijay Janapa Reddi

As machine learning shifts from laboratory curiosity to critical infrastructure, the systems that sustain it span an extraordinary range, from sub-milliwatt microcontrollers to mul…

cs.DC2026

MLCommons Chakra: Advancing Performance Benchmarking and Co-design using Standardized Execution Traces

Srinivas Sridharan, Theodor-Adrian Badea, Andy Balogh +26

The fast pace of artificial intelligence~(AI) innovation demands an agile methodology for observation, reproduction and optimization of distributed machine learning~(ML) workload b…

cs.DC2025

SwizzlePerf: Hardware-Aware LLMs for GPU Kernel Performance Optimization

Arya Tschand, Muhammad Awad, Ryan Swann +5

Large language models (LLMs) have shown progress in GPU kernel performance engineering using inefficient search-based methods that optimize around runtime. Any existing approach la…

cs.DC20251 cited

COSMIC: Enabling Full-Stack Co-Design and Optimization of Distributed Machine Learning Systems

Aditi Raju, Jared Ni, William Won +6

Large-scale machine learning models necessitate distributed systems, posing significant design challenges due to the large parameter space across distinct design stacks. Existing s…

cs.DC20205 cited

AI Tax: The Hidden Cost of AI Data Center Applications

Daniel Richins, Dharmisha Doshi, Matthew Blackmore +9

Artificial intelligence and machine learning are experiencing widespread adoption in industry and academia. This has been driven by rapid advances in the applications and accuracy…