20 citations · 39 across the 3 of their papers we have counts for
3 papers
cs.AR2021★ 19 cited
First-Generation Inference Accelerator Deployment at Facebook
Michael Anderson, Benny Chen, Stephen Chen +112
In this paper, we provide a deep dive into the deployment of inference accelerators at Facebook. Many of our ML workloads have unique characteristics, such as sparse memory accesse…
cs.LG2021
Low-Precision Hardware Architectures Meet Recommendation Model Inference at Scale
Zhaoxia, Deng, Jongsoo Park +17
Tremendous success of machine learning (ML) and the unabated growth in ML model complexity motivated many ML-specific designs in both CPU and accelerator architectures to speed up…
cs.LG2021★ 20 cited
FBGEMM: Enabling High-Performance Low-Precision Deep Learning Inference
Daya Khudia, Jianyu Huang, Protonu Basu +4
Deep learning models typically use single-precision (FP32) floating point data types for representing activations and weights, but a slew of recent research work has shown that com…