1 citations · 1 across the 1 of their papers we have counts for
3 papers
cs.DC2025★ 1 cited
TokenScale: Timely and Accurate Autoscaling for Disaggregated LLM Serving with Token Velocity
Ruiqi Lai, Hongrui Liu, Chengzhi Lu +6
The architectural shift to prefill/decode (PD) disaggregation in LLM serving improves resource utilization but struggles with the bursty nature of modern workloads. Existing autosc…
cs.AR2025
BitDecoding: Unlocking Tensor Cores for Long-Context LLMs with Low-Bit KV Cache
Dayou Du, Shijie Cao, Jianyi Cheng +3
The growth of long-context Large Language Models (LLMs) significantly increases memory and bandwidth pressure during autoregressive decoding due to the expanding Key-Value (KV) cac…
cs.LG2025
WaferLLM: Large Language Model Inference at Wafer Scale
Congjie He, Yeqi Huang, Pei Mu +5
Emerging AI accelerators increasingly adopt wafer-scale manufacturing technologies, integrating hundreds of thousands of AI cores in a mesh architecture with large distributed on-c…