26 citations · 28 across the 3 of their papers we have counts for
3 papers
Binarized Neural Machine Translation
Yichi Zhang, Ankush Garg, Yuan Cao +4
The rapid scaling of language models is motivating research using low-bitwidth quantization. In this work, we propose a novel binarization technique for Transformers applied to mac…
4-bit Conformer with Native Quantization Aware Training for Speech Recognition
Shaojin Ding, Phoenix Meadowlark, Yanzhang He +3
Reducing the latency and model size has always been a significant research problem for live Automatic Speech Recognition (ASR) application scenarios. Along this direction, model qu…
Pareto-Optimal Quantized ResNet Is Mostly 4-bit
AmirAli Abdolrashidi, Lisa Wang, Shivani Agrawal +4
Quantization has become a popular technique to compress neural networks and reduce compute cost, but most prior work focuses on studying quantization without changing the network s…