Publications (19)
DataPerf: Benchmarks for Data-Centric AI Development
Mark Mazumder, Colby Banbury, Xiaozhe Yao +42
Machine learning research has long focused on models rather than datasets, and prominent datasets are used for common ML tasks without regard to the breadth, difficulty, and faithf…
Data Engineering for Everyone
Vijay Janapa Reddi, Greg Diamos, Pete Warden +2
Data engineering is one of the fastest-growing fields within machine learning (ML). As ML becomes more common, the appetite for data grows more ravenous. But ML requires more data…
MLPerf Training Benchmark
Peter Mattson, Christine Cheng, Cody Coleman +34
Machine learning (ML) needs industry-standard performance benchmarks to support design and competitive evaluation of the many emerging software and hardware solutions for ML. But M…
AILuminate: Introducing v1.0 of the AI Risk and Reliability Benchmark from MLCommons
Shaona Ghosh, Heather Frase, Adina Williams +99
The rapid advancement and deployment of AI systems have created an urgent need for standard safety-evaluation frameworks. This paper introduces AILuminate v1.0, the first comprehen…
A Technical Policy Blueprint for Trustworthy Decentralized AI
Hasan Kassem, Orion Banks, Omar Benjelloun +20
Decentralized AI systems, such as federated learning, can play a critical role in further unlocking AI asset marketplaces (e.g., healthcare data marketplaces) thanks to increased a…
MLPerf Inference Benchmark
Vijay Janapa Reddi, Christine Cheng, David Kanter +44
Machine-learning (ML) hardware and software system demand is burgeoning. Driven by ML applications, the number of different ML inference systems has exploded. Over 100 organization…
GaNDLF: A Generally Nuanced Deep Learning Framework for Scalable End-to-End Clinical Workflows in Medical Imaging
Sarthak Pati, Siddhesh P. Thakur, İbrahim Ethem Hamamcı +39
Deep Learning (DL) has the potential to optimize machine learning in both the scientific and clinical communities. However, greater expertise is required to develop DL algorithms,…
Scale MLPerf-0.6 models on Google TPU-v3 Pods
Sameer Kumar, Victor Bitorff, Dehao Chen +9
The recent submission of Google TPU-v3 Pods to the industry wide MLPerf v0.6 training benchmark demonstrates the scalability of a suite of industry relevant ML models. MLPerf defin…
DMLR: Data-centric Machine Learning Research -- Past, Present and Future
Luis Oala, Manil Maskey, Lilith Bat-Leah +35
Drawing from discussions at the inaugural DMLR workshop at ICML 2023 and meetings prior, in this report we outline the relevance of community engagement and infrastructure developm…
MLPerf HPC: A Holistic Benchmark Suite for Scientific Machine Learning on HPC Systems
Steven Farrell, Murali Emani, Jacob Balma +40
Scientific communities are increasingly adopting machine learning and deep learning models in their applications to accelerate scientific insights. High performance computing syste…
Dynatask: A Framework for Creating Dynamic AI Benchmark Tasks
Tristan Thrush, Kushal Tirumala, Anmol Gupta +7
We introduce Dynatask: an open source system for setting up custom NLP tasks that aims to greatly lower the technical knowledge and effort required for hosting and evaluating state…
Common Limitations of Image Processing Metrics: A Picture Story
Annika Reinke, Minu D. Tizabi, Carole H. Sudre +90
While the importance of automatic image analysis is continuously increasing, recent meta-research revealed major flaws with respect to algorithm validation. Performance metrics are…
Introducing v0.5 of the AI Safety Benchmark from MLCommons
Bertie Vidgen, Adarsh Agrawal, Ahmed M. Ahmed +97
This paper introduces v0.5 of the AI Safety Benchmark, which has been created by the MLCommons AI Safety Working Group. The AI Safety Benchmark has been designed to assess the safe…
Understanding metric-related pitfalls in image analysis validation
Annika Reinke, Minu D. Tizabi, Michael Baumgartner +75
Validation metrics are key for the reliable tracking of scientific progress and for bridging the current chasm between artificial intelligence (AI) research and its translation int…
Croissant: A Metadata Format for ML-Ready Datasets
Mubashara Akhtar, Omar Benjelloun, Costanza Conforti +28
Data is a critical resource for machine learning (ML), yet working with data remains a key friction point. This paper introduces Croissant, a metadata format for datasets that crea…
Metrics reloaded: Recommendations for image analysis validation
Lena Maier-Hein, Annika Reinke, Patrick Godau +71
Increasing evidence shows that flaws in machine learning (ML) algorithm validation are an underestimated global problem. Particularly in automatic biomedical image analysis, chosen…
Benchmarking Neural Network Training Algorithms
George E. Dahl, Frank Schneider, Zachary Nado +22
Training algorithms, broadly construed, are an essential part of every deep learning pipeline. Training algorithm improvements that speed up training across a wide variety of workl…
MLPerf Mobile Inference Benchmark
Vijay Janapa Reddi, David Kanter, Peter Mattson +24
This paper presents the first industry-standard open-source machine learning (ML) benchmark to allow perfor mance and accuracy evaluation of mobile devices with different AI chips…
MedPerf: Open Benchmarking Platform for Medical Artificial Intelligence using Federated Evaluation
Alexandros Karargyris, Renato Umeton, Micah J. Sheller +39
Medical AI has tremendous potential to advance healthcare by supporting the evidence-based practice of medicine, personalizing patient treatment, reducing costs, and improving prov…