4 papers
Power Aware Dynamic Reallocation For Inference
Yiwei Jiang, Sangeeta Chowdhary, Nathaniel Morris +3
Disaggregation has emerged as a powerful strategy for optimizing large language model (LLM) inference by separating compute-intensive prefill and memory-bound decode phases across…
AI Benchmark Democratization and Carpentry
Gregor von Laszewski, Wesley Brewer, Jeyan Thiyagalingam +28
Benchmarks are a cornerstone of modern machine learning, enabling reproducibility, comparison, and scientific progress. However, AI benchmarks are increasingly complex, requiring d…
An MLCommons Scientific Benchmarks Ontology
Ben Hawks, Gregor von Laszewski, Matthew D. Sinclair +6
Scientific machine learning research spans diverse domains and data modalities, yet existing benchmark efforts remain siloed and lack standardization. This makes novel and transfor…
PAL: A Variability-Aware Policy for Scheduling ML Workloads in GPU Clusters
Rutwik Jain, Brandon Tran, Keting Chen +2
Large-scale computing systems are increasingly using accelerators such as GPUs to enable peta- and exa-scale levels of compute to meet the needs of Machine Learning (ML) and scient…