571 citations · 718 across the 4 of their papers we have counts for
4 papers
Low-Precision Hardware Architectures Meet Recommendation Model Inference at Scale
Zhaoxia, Deng, Jongsoo Park +17
Tremendous success of machine learning (ML) and the unabated growth in ML model complexity motivated many ML-specific designs in both CPU and accelerator architectures to speed up…
Practical optimization for hybrid quantum-classical algorithms
Gian Giacomo Guerreschi, Mikhail Smelyanskiy
A novel class of hybrid quantum-classical algorithms based on the variational approach have recently emerged from separate proposals addressing, for example, quantum chemistry and…
On Large-Batch Training for Deep Learning: Generalization Gap and Sharp Minima
Nitish Shirish Keskar, Dheevatsa Mudigere, Jorge Nocedal +2
The stochastic gradient descent (SGD) method and its variants are algorithms of choice for many Deep Learning tasks. These methods operate in a small-batch regime wherein a fractio…
Lattice QCD with Domain Decomposition on Intel Xeon Phi Co-Processors
Simon Heybrock, Bálint Joó, Dhiraj D. Kalamkar +4
The gap between the cost of moving data and the cost of computing continues to grow, making it ever harder to design iterative solvers on extreme-scale architectures. This problem…