efficiency analysis 1kv cache sparsification 1large language models 1long-context inference 1speculative decoding 1
From the 1 of 3 linked papers with an AI index.
3 papers
cs.DC2026
Energy-Efficient LLM Serving via Disaggregated Attention--FFN and Flexible Frequency Scaling
Cunchen Hu, Liangliang Xu, Tian Liu +9
Large language model (LLM) serving spans diverse applications with stringent service-level objectives (SLOs), often requiring GPUs to run at maximum frequencies and increasing ener…
cs.CL2026
A Sparse Glimpse of the Whole: Train-Free Self-Speculative Decoding
Yuesong Liu, Yuan Zeng, Min Lyu +3
The paper proposes SparseSpec-L, a training-free self-speculative decoding method that uses a sparsified key‑value cache and an entropy‑based controller to speed up long‑context in…
cs.DC2025
New Wide Locally Recoverable Codes with Unified Locality
Liangliang Xu, Fengming Tang, Tingting Chen +3
Wide Locally Recoverable Codes (LRCs) have recently been proposed as a solution for achieving high reliability, good performance, and ultra-low storage cost in distributed storage…