Showing cs.CLShow all
3 papers · 1 filter
cs.CL2025
LoopServe: An Adaptive Dual-phase LLM Inference Acceleration System for Multi-Turn Dialogues
Haoyang Li, Zhanchao Xu, Yiming Li +9
Multi-turn dialogues are essential in many real-world applications of large language models, such as chatbots and virtual assistants. As conversation histories become longer, exist…
cs.CL2025
GTA: Grouped-head latenT Attention
Luoyang Sun, Cheng Deng, Jiwen Jiang +5
Attention mechanisms underpin the success of large language models (LLMs), yet their substantial computational and memory overhead poses challenges for optimizing efficiency and pe…
cs.CL2025
Activation-aware Probe-Query: Effective Key-Value Retrieval for Long-Context LLMs Inference
Qingfa Xiao, Jiachuan Wang, Haoyang Li +6
Recent advances in large language models (LLMs) have showcased exceptional performance in long-context tasks, while facing significant inference efficiency challenges with limited…