2 citations · 2 across the 2 of their papers we have counts for
2 papers
cs.LG2026★ 2 cited
Towards Efficient Large Language Model Serving: A Survey on System-Aware KV Cache Optimization
Jiantong Jiang, Peiyu Yang, Rui Zhang +1
Despite the rapid advancements of large language models (LLMs), LLM serving systems remain memory-intensive and costly. The key-value (KV) cache, which stores KV tensors during aut…
cs.SD2026
CleanCodec: Efficient and Robust Speech Tokenization via Perceptually Guided Encoding
Eugene Kwek, Feng Liu, Rui Zhang +1
Neural audio codecs are a key component of speech processing pipelines, compressing audio into discrete tokens for downstream modeling. However, existing codecs struggle to balance…