2 papers
cs.CL2026
Self-Speculative Biased Decoding for Faster Re-Translation
Linxiao Zeng, Haoyun Deng, Kangyuan Shu +1
Large language models achieve strong machine translation quality but incur high inference cost and latency, posing challenges for simultaneous translation. Re-translation provides…
cs.DC2025
Enabling Dynamic Sparsity in Quantized LLM Inference
Rongxiang Wang, Kangyuan Shu, Felix Xiaozhu Lin
Deploying large language models (LLMs) on end-user devices is gaining importance due to benefits in responsiveness, privacy, and operational cost. Yet the limited memory and comput…