1 citations · 1 across the 2 of their papers we have counts for
2 papers
cs.CL2026
DASH-KV: Accelerating Long-Context LLM Inference via Asymmetric KV Cache Hashing
Jinyu Guo, Zhihan Zhang, Jiehui Xie +7
The quadratic computational complexity of the standard attention mechanism constitutes a fundamental bottleneck for large language models in long-context inference. While existing…
cs.AR2024★ 1 cited
LlamaF: An Efficient Llama2 Architecture Accelerator on Embedded FPGAs
Han Xu, Yutong Li, Shihao Ji
Large language models (LLMs) have demonstrated remarkable abilities in natural language processing. However, their deployment on resource-constrained embedded devices remains diffi…