activity
20172021
most citedMLPerf Training Benchmark

171 citations · 257 across the 6 of their papers we have counts for

collaborators

8 papers

cs.DC202038 cited

DeepRecSys: A System for Optimizing End-To-End At-scale Neural Recommendation Inference

Udit Gupta, Samuel Hsia, Vikram Saraph +6

Neural personalized recommendation is the corner-stone of a wide collection of cloud services and products, constituting significant compute demand of the cloud infrastructure. Thu…

cs.AR2020

CHIPKIT: An agile, reusable open-source framework for rapid test chip development

Paul Whatmough, Marco Donato, Glenn Ko +3

The current trend for domain-specific architectures (DSAs) has led to renewed interest in research test chips to demonstrate new specialized hardware. Tape-outs also offer huge ped…

cs.LG201910 cited

SMAUG: End-to-End Full-Stack Simulation Infrastructure for Deep Learning Workloads

Sam Likun Xi, Yuan Yao, Kshitij Bhardwaj +3

In recent years, there has been tremendous advances in hardware acceleration of deep neural networks. However, most of the research has focused on optimizing accelerator microarchi…

cs.LG2019171 cited

MLPerf Training Benchmark

Peter Mattson, Christine Cheng, Cody Coleman +34

Machine learning (ML) needs industry-standard performance benchmarks to support design and competitive evaluation of the many emerging software and hardware solutions for ML. But M…

cs.LG2019

AdaptivFloat: A Floating-point based Data Type for Resilient Deep Learning Inference

Thierry Tambe, En-Yu Yang, Zishen Wan +5

Conventional hardware-friendly quantization methods, such as fixed-point or integer, tend to perform poorly at very low word sizes as their shrinking dynamic ranges cannot adequate…

eess.SP2019

MASR: A Modular Accelerator for Sparse RNNs

Udit Gupta, Brandon Reagen, Lillian Pentecost +5

Recurrent neural networks (RNNs) are becoming the de facto solution for speech recognition. RNNs exploit long-term temporal relationships in data by applying repeated, learned tran…