3 papers
cs.LG2025
OmniDraft: A Cross-vocabulary, Online Adaptive Drafter for On-device Speculative Decoding
Ramchalam Kinattinkara Ramakrishnan, Zhaocong Yuan, Shaojie Zhuo +4
Speculative decoding generally dictates having a small, efficient draft model that is either pretrained or distilled offline to a particular target model series, for instance, Llam…
cs.SD2025
Edge-ASR: Towards Low-Bit Quantization of Automatic Speech Recognition Models
Chen Feng, Yicheng Lin, Shaojie Zhuo +4
Recent advances in Automatic Speech Recognition (ASR) have demonstrated remarkable accuracy and robustness in diverse audio applications, such as live transcription and voice comma…
cs.LG2024
Stepping Forward on the Last Mile
Chen Feng, Shaojie Zhuo, Xiaopeng Zhang +3
Continuously adapting pre-trained models to local data on resource constrained edge devices is the for model deployment. However, as models increase in size and…