3 papers
cs.DC2025
Enabling Dynamic Sparsity in Quantized LLM Inference
Rongxiang Wang, Kangyuan Shu, Felix Xiaozhu Lin
Deploying large language models (LLMs) on end-user devices is gaining importance due to benefits in responsiveness, privacy, and operational cost. Yet the limited memory and comput…
cs.CL2025
Self-Speculative Biased Decoding for Faster Re-Translation
Linxiao Zeng, Haoyun Deng, Kangyuan Shu +1
Large language models achieve strong machine translation quality but incur high inference cost and latency, posing challenges for simultaneous translation. Re-translation provides…
cs.CL2022
Streaming, fast and accurate on-device Inverse Text Normalization for Automatic Speech Recognition
Yashesh Gaur, Nick Kibre, Jian Xue +5
Automatic Speech Recognition (ASR) systems typically yield output in lexical form. However, humans prefer a written form output. To bridge this gap, ASR systems usually employ Inve…