2 citations · 3 across the 3 of their papers we have counts for
3 papers
MRQ:Support Multiple Quantization Schemes through Model Re-Quantization
Manasa Manohara, Sankalp Dayal, Tariq Afzal +2
Despite the proliferation of diverse hardware accelerators (e.g., NPU, TPU, DPU), deploying deep learning models on edge devices with fixed-point hardware is still challenging due…
Accelerator-Aware Training for Transducer-Based Speech Recognition
Suhaila M. Shakiah, Rupak Vignesh Swaminathan, Hieu Duy Nguyen +6
Machine learning model weights and activations are represented in full-precision during training. This leads to performance degradation in runtime when deployed on neural network a…
Sub-8-Bit Quantization Aware Training for 8-Bit Neural Network Accelerator with On-Device Speech Recognition
Kai Zhen, Hieu Duy Nguyen, Raviteja Chinta +4
We present a novel sub-8-bit quantization-aware training (S8BQAT) scheme for 8-bit neural network accelerators. Our method is inspired from Lloyd-Max compression theory with practi…