4 citations · 4 across the 2 of their papers we have counts for
3 papers
cs.CL2022
Zero-Shot Dynamic Quantization for Transformer Inference
Yousef El-Kurdi, Jerry Quinn, Avirup Sil
We introduce a novel run-time method for significantly reducing the accuracy loss associated with quantizing BERT-like models to 8-bit integers. Existing methods for quantizing mod…
cs.LG2019★ 4 cited
Optimal Mini-Batch Size Selection for Fast Gradient Descent
Michael P. Perrone, Haidar Khan, Changhoan Kim +3
This paper presents a methodology for selecting the mini-batch size that minimizes Stochastic Gradient Descent (SGD) learning time for single and multiple learner problems. By deco…
cs.CL2018
Pieces of Eight: 8-bit Neural Machine Translation
Jerry Quinn, Miguel Ballesteros
Neural machine translation has achieved levels of fluency and adequacy that would have been surprising a short time ago. Output quality is extremely relevant for industry purposes,…