2 papers
cs.CL2026
HyLRA: Hybrid Layer Reuse Attention for Efficient Long-Context Inference
Xuan Ai, Qingqing Yang, Peng Wang +4
Long-context inference in Large Language Models (LLMs) is bottlenecked by the quadratic computation complexity of attention and the substantial memory footprint of Key-Value (KV) c…
cs.LG2025
A Mathematical Theory of Top- Sparse Attention via Total Variation Distance
Georgios Tzachristas, Lei Deng, Ioannis Tzachristas +2
We develop a unified mathematical framework for certified Top- attention truncation that quantifies approximation error at both the distribution and output levels. For a single…