activity
20172022
most citedMLPerf Training Benchmark

171 citations · 272 across the 11 of their papers we have counts for

collaborators
Showing cs.LGShow all

6 papers · 1 filter

cs.LG201910 cited

SMAUG: End-to-End Full-Stack Simulation Infrastructure for Deep Learning Workloads

Sam Likun Xi, Yuan Yao, Kshitij Bhardwaj +3

In recent years, there has been tremendous advances in hardware acceleration of deep neural networks. However, most of the research has focused on optimizing accelerator microarchi…

cs.LG2019

A binary-activation, multi-level weight RNN and training algorithm for ADC-/DAC-free and noise-resilient processing-in-memory inference with eNVM

Siming Ma, David Brooks, Gu-Yeon Wei

We propose a new algorithm for training neural networks with binary activations and multi-level weights, which enables efficient processing-in-memory circuits with embedded nonvola…

cs.LG2019171 cited

MLPerf Training Benchmark

Peter Mattson, Christine Cheng, Cody Coleman +34

Machine learning (ML) needs industry-standard performance benchmarks to support design and competitive evaluation of the many emerging software and hardware solutions for ML. But M…

cs.LG2019

AdaptivFloat: A Floating-point based Data Type for Resilient Deep Learning Inference

Thierry Tambe, En-Yu Yang, Zishen Wan +5

Conventional hardware-friendly quantization methods, such as fixed-point or integer, tend to perform poorly at very low word sizes as their shrinking dynamic ranges cannot adequate…

cs.LG2019

Benchmarking TPU, GPU, and CPU Platforms for Deep Learning

Yu Emma Wang, Gu-Yeon Wei, David Brooks

Training deep learning models is compute-intensive and there is an industry-wide trend towards hardware specialization to improve performance. To systematically benchmark deep lear…

cs.LG201728 cited

Weightless: Lossy Weight Encoding For Deep Neural Network Compression

Brandon Reagen, Udit Gupta, Robert Adolf +4

The large memory requirements of deep neural networks limit their deployment and adoption on many devices. Model compression methods effectively reduce the memory requirements of t…