19 citations · 28 across the 2 of their papers we have counts for
4 papers
First-Generation Inference Accelerator Deployment at Facebook
Michael Anderson, Benny Chen, Stephen Chen +112
In this paper, we provide a deep dive into the deployment of inference accelerators at Facebook. Many of our ML workloads have unique characteristics, such as sparse memory accesse…
Alternate Model Growth and Pruning for Efficient Training of Recommendation Systems
Xiaocong Du, Bhargav Bhushanam, Jiecao Yu +7
Deep learning recommendation systems at scale have provided remarkable gains through increasing model capacity (i.e. wider and deeper neural networks), but it comes at significant…
Adaptive Dense-to-Sparse Paradigm for Pruning Online Recommendation System with Non-Stationary Data
Mao Ye, Dhruv Choudhary, Jiecao Yu +6
Large scale deep learning provides a tremendous opportunity to improve the quality of content recommendation systems by employing both wider and deeper models, but this comes at gr…
Spatial-Winograd Pruning Enabling Sparse Winograd Convolution
Jiecao Yu, Jongsoo Park, Maxim Naumov
Deep convolutional neural networks (CNNs) are deployed in various applications but demand immense computational requirements. Pruning techniques and Winograd convolution are two ty…