171 citations · 257 across the 6 of their papers we have counts for
8 papers
DeepRecSys: A System for Optimizing End-To-End At-scale Neural Recommendation Inference
Udit Gupta, Samuel Hsia, Vikram Saraph +6
Neural personalized recommendation is the corner-stone of a wide collection of cloud services and products, constituting significant compute demand of the cloud infrastructure. Thu…
CHIPKIT: An agile, reusable open-source framework for rapid test chip development
Paul Whatmough, Marco Donato, Glenn Ko +3
The current trend for domain-specific architectures (DSAs) has led to renewed interest in research test chips to demonstrate new specialized hardware. Tape-outs also offer huge ped…
SMAUG: End-to-End Full-Stack Simulation Infrastructure for Deep Learning Workloads
Sam Likun Xi, Yuan Yao, Kshitij Bhardwaj +3
In recent years, there has been tremendous advances in hardware acceleration of deep neural networks. However, most of the research has focused on optimizing accelerator microarchi…
MLPerf Training Benchmark
Peter Mattson, Christine Cheng, Cody Coleman +34
Machine learning (ML) needs industry-standard performance benchmarks to support design and competitive evaluation of the many emerging software and hardware solutions for ML. But M…
AdaptivFloat: A Floating-point based Data Type for Resilient Deep Learning Inference
Thierry Tambe, En-Yu Yang, Zishen Wan +5
Conventional hardware-friendly quantization methods, such as fixed-point or integer, tend to perform poorly at very low word sizes as their shrinking dynamic ranges cannot adequate…
MASR: A Modular Accelerator for Sparse RNNs
Udit Gupta, Brandon Reagen, Lillian Pentecost +5
Recurrent neural networks (RNNs) are becoming the de facto solution for speech recognition. RNNs exploit long-term temporal relationships in data by applying repeated, learned tran…