most citedActivation-aware Probe-Query: Effective Key-Value Retrieval for Long-Context LLMs Inference

1 citations · 1 across the 3 of their papers we have counts for

collaborators

5 papers

cs.CL2025

GTA: Grouped-head latenT Attention

Luoyang Sun, Cheng Deng, Jiwen Jiang +5

Attention mechanisms underpin the success of large language models (LLMs), yet their substantial computational and memory overhead poses challenges for optimizing efficiency and pe…

cs.CL2025

LoopServe: An Adaptive Dual-phase LLM Inference Acceleration System for Multi-Turn Dialogues

Haoyang Li, Zhanchao Xu, Yiming Li +9

Multi-turn dialogues are essential in many real-world applications of large language models, such as chatbots and virtual assistants. As conversation histories become longer, exist…

cs.MA2025

MF-LLM: Simulating Population Decision Dynamics via a Mean-Field Large Language Model Framework

Qirui Mi, Mengyue Yang, Xiangning Yu +6

Simulating collective decision-making involves more than aggregating individual behaviors; it emerges from dynamic interactions among individuals. While large language models (LLMs…

cs.CL2025

PLM: Efficient Peripheral Language Models Hardware-Co-Designed for Ubiquitous Computing

Cheng Deng, Luoyang Sun, Jiwen Jiang +10

While scaling laws have been continuously validated in large language models (LLMs) with increasing model parameters, the inherent tension between the inference demands of LLMs and…

cs.CL20251 cited

Activation-aware Probe-Query: Effective Key-Value Retrieval for Long-Context LLMs Inference

Qingfa Xiao, Jiachuan Wang, Haoyang Li +6

Recent advances in large language models (LLMs) have showcased exceptional performance in long-context tasks, while facing significant inference efficiency challenges with limited…