Showing cs.CLShow all
2 papers · 1 filter
cs.CL2025
Rhea: Role-aware Heuristic Episodic Attention for Conversational LLMs
Wanyang Hong, Zhaoning Zhang, Yi Chen +5
Large Language Models (LLMs) have achieved remarkable performance on single-turn tasks, yet their effectiveness deteriorates in multi-turn conversations. We define this phenomenon…
cs.CL2025
Dovetail: A CPU/GPU Heterogeneous Speculative Decoding for LLM inference
Libo Zhang, Zhaoning Zhang, Baizhou Xu +4
With the continuous advancement in the performance of large language models (LLMs), their demand for computational resources and memory has significantly increased, which poses maj…