3 papers
eess.SP2026
AVSG: Accelerated Vectorized Sparse Gather for Efficient KV Cache Offload in Sparse-Attention LLM Serving
Wenwei Kuang, Xiangyu Wang, Chong Wu +5
Dynamic sparse attention reduces long-context attention computation by selecting only a subset of tokens, but still requires access to the full KV cache, leaving serving memory-bou…
cs.DC2025
OmniInfer: System-Wide Acceleration Techniques for Optimizing LLM Serving Throughput and Latency
Jun Wang, Yunxiang Yao, Wenwei Kuang +11
Large Language Models drive a wide range of modern AI applications but impose substantial challenges on large-scale serving systems due to intensive computation, strict latency con…
cs.LG2024
A Mean Field Ansatz for Zero-Shot Weight Transfer
Xingyuan Chen, Wenwei Kuang, Lei Deng +3
The pre-training cost of large language models (LLMs) is prohibitive. One cutting-edge approach to reduce the cost is zero-shot weight transfer, also known as model growth for some…