19 citations · 19 across the 2 of their papers we have counts for
2 papers
cs.AR2021
Supporting Massive DLRM Inference Through Software Defined Memory
Ehsan K. Ardestani, Changkyu Kim, Seung Jae Lee +17
Deep Learning Recommendation Models (DLRM) are widespread, account for a considerable data center footprint, and grow by more than 1.5x per year. With model size soon to be in tera…
cs.AR2021★ 19 cited
First-Generation Inference Accelerator Deployment at Facebook
Michael Anderson, Benny Chen, Stephen Chen +112
In this paper, we provide a deep dive into the deployment of inference accelerators at Facebook. Many of our ML workloads have unique characteristics, such as sparse memory accesse…