3 papers
cs.DC2026
OasisKV: Scaling In-Decode KV Cache Beyond HBM with Lookahead Sparse Prefetching
Can Xiao, Sukmin Cho, Junbong We +7
Large language model (LLM) inference serving is increasingly constrained by memory rather than compute. As long-context and long-form reasoning workloads become more prevalent, the…
cs.LG2026
An Efficient Hybrid Sparse Attention with CPU-GPU Parallelism for Long-Context Inference
Feiyu Yao, Zhixiong Niu, Xiaqing Li +3
Long-context inference increasingly operates over CPU-resident KV caches, either because decoding-time KV states exceed GPU memory capacity or because disaggregated prefill-decode…
cs.NI2025
Automating Conflict-Aware ACL Configurations with Natural Language Intents
Wenlong Ding, Jianqiang Li, Zhixiong Niu +3
ACL configuration is essential for managing network flow reachability, yet its complexity grows significantly with topologies and pre-existing rules. To carry out ACL configuration…