4 papers
Energy-Efficient LLM Serving via Disaggregated Attention--FFN and Flexible Frequency Scaling
Cunchen Hu, Liangliang Xu, Tian Liu +9
Large language model (LLM) serving spans diverse applications with stringent service-level objectives (SLOs), often requiring GPUs to run at maximum frequencies and increasing ener…
A Sparse Glimpse of the Whole: Train-Free Self-Speculative Decoding
Yuesong Liu, Yuan Zeng, Min Lyu +3
Speculative decoding alleviates the memory-bandwidth bottleneck in large language model inference, but its acceleration is jointly constrained by drafting overhead, token acceptanc…
New Wide Locally Recoverable Codes with Unified Locality
Liangliang Xu, Fengming Tang, Tingting Chen +3
Wide Locally Recoverable Codes (LRCs) have recently been proposed as a solution for achieving high reliability, good performance, and ultra-low storage cost in distributed storage…
Deterministic Data Distribution for Efficient Recovery in Erasure-Coded Storage Systems
Liangliang Xu, Min Lyu, Zhipeng Li +2
Due to individual unreliable commodity components, failures are common in large-scale distributed storage systems. Erasure codes are widely deployed in practical storage systems to…