Showing cs.CLShow all
2 papers · 1 filter
cs.CL2025
Speculative Decoding Meets Quantization: Compatibility Evaluation and Hierarchical Framework Design
Yudi Zhang, Weilin Zhao, Xu Han +4
Speculative decoding and quantization effectively accelerate memory-bound inference of large language models. Speculative decoding mitigates the memory bandwidth bottleneck by veri…
cs.CL2024
Enabling Real-Time Conversations with Minimal Training Costs
Wang Xu, Shuo Wang, Weilin Zhao +6
Large language models (LLMs) have demonstrated the ability to improve human efficiency through conversational interactions. Conventional LLM-powered dialogue systems, operating on…