activity
20172023
most citedMLPerf Training Benchmark

171 citations · 296 across the 16 of their papers we have counts for

collaborators
Showing 2019Show all

9 papers · 1 filter

cs.DC20197 cited

RecNMP: Accelerating Personalized Recommendation with Near-Memory Processing

Liu Ke, Udit Gupta, Carole-Jean Wu +18

Personalized recommendation systems leverage deep learning models and account for the majority of data center AI cycles. Their performance is dominated by memory-bound sparse embed…

cs.LG201910 cited

SMAUG: End-to-End Full-Stack Simulation Infrastructure for Deep Learning Workloads

Sam Likun Xi, Yuan Yao, Kshitij Bhardwaj +3

In recent years, there has been tremendous advances in hardware acceleration of deep neural networks. However, most of the research has focused on optimizing accelerator microarchi…

cs.LG2019

A binary-activation, multi-level weight RNN and training algorithm for ADC-/DAC-free and noise-resilient processing-in-memory inference with eNVM

Siming Ma, David Brooks, Gu-Yeon Wei

We propose a new algorithm for training neural networks with binary activations and multi-level weights, which enables efficient processing-in-memory circuits with embedded nonvola…

cs.LG2019171 cited

MLPerf Training Benchmark

Peter Mattson, Christine Cheng, Cody Coleman +34

Machine learning (ML) needs industry-standard performance benchmarks to support design and competitive evaluation of the many emerging software and hardware solutions for ML. But M…

cs.LG2019

AdaptivFloat: A Floating-point based Data Type for Resilient Deep Learning Inference

Thierry Tambe, En-Yu Yang, Zishen Wan +5

Conventional hardware-friendly quantization methods, such as fixed-point or integer, tend to perform poorly at very low word sizes as their shrinking dynamic ranges cannot adequate…

eess.SP2019

MASR: A Modular Accelerator for Sparse RNNs

Udit Gupta, Brandon Reagen, Lillian Pentecost +5

Recurrent neural networks (RNNs) are becoming the de facto solution for speech recognition. RNNs exploit long-term temporal relationships in data by applying repeated, learned tran…