1 citations · 1 across the 3 of their papers we have counts for
4 papers · 1 filter
GTA: Grouped-head latenT Attention
Luoyang Sun, Cheng Deng, Jiwen Jiang +5
Attention mechanisms underpin the success of large language models (LLMs), yet their substantial computational and memory overhead poses challenges for optimizing efficiency and pe…
LoopServe: An Adaptive Dual-phase LLM Inference Acceleration System for Multi-Turn Dialogues
Haoyang Li, Zhanchao Xu, Yiming Li +9
Multi-turn dialogues are essential in many real-world applications of large language models, such as chatbots and virtual assistants. As conversation histories become longer, exist…
PLM: Efficient Peripheral Language Models Hardware-Co-Designed for Ubiquitous Computing
Cheng Deng, Luoyang Sun, Jiwen Jiang +10
While scaling laws have been continuously validated in large language models (LLMs) with increasing model parameters, the inherent tension between the inference demands of LLMs and…
Activation-aware Probe-Query: Effective Key-Value Retrieval for Long-Context LLMs Inference
Qingfa Xiao, Jiachuan Wang, Haoyang Li +6
Recent advances in large language models (LLMs) have showcased exceptional performance in long-context tasks, while facing significant inference efficiency challenges with limited…