184 citations · 189 across the 3 of their papers we have counts for
3 papers
Augmenting Hessians with Inter-Layer Dependencies for Mixed-Precision Post-Training Quantization
Clemens JS Schaefer, Navid Lambert-Shirzad, Xiaofan Zhang +7
Efficiently serving neural network models with low latency is becoming more challenging due to increasing model complexity and parameter count. Model quantization offers a solution…
Mixed Precision Post Training Quantization of Neural Networks with Sensitivity Guided Search
Clemens JS Schaefer, Elfie Guo, Caitlin Stanton +7
Serving large-scale machine learning (ML) models efficiently and with low latency has become challenging owing to increasing model size and complexity. Quantizing models can simult…
Lingvo: a Modular and Scalable Framework for Sequence-to-Sequence Modeling
Jonathan Shen, Patrick Nguyen, Yonghui Wu +88
Lingvo is a Tensorflow framework offering a complete solution for collaborative deep learning research, with a particular focus towards sequence-to-sequence models. Lingvo models a…