67 citations · 90 across the 16 of their papers we have counts for
Showing cs.CLShow all
2 papers · 1 filter
cs.CL2023
TODM: Train Once Deploy Many Efficient Supernet-Based RNN-T Compression For On-device ASR Models
Yuan Shangguan, Haichuan Yang, Danni Li +11
Automatic Speech Recognition (ASR) models need to be optimized for specific hardware before they can be deployed on devices. This can be done by tuning the model's hyperparameters…
cs.CL2023★ 15 cited
LLM-QAT: Data-Free Quantization Aware Training for Large Language Models
Zechun Liu, Barlas Oguz, Changsheng Zhao +6
Several post-training quantization methods have been applied to large language models (LLMs), and have been shown to perform well down to 8-bits. We find that these methods break d…