3 papers
cs.NI2026
An Internet for the KV Cache: Rethinking Classical Infrastructure Boundaries in the LLM Inference Age
Siddhant Ray, Nick Feamster, Junchen Jiang
LLM inference has become a global-scale, heterogeneous workload spanning agents, retrieval, tool-use, code execution and multi-modal reasoning. These workloads naturally enable con…
cs.NI2026
SwiftQueue: Optimizing Low-Latency Applications with Swift Packet Queuing
Siddhant Ray, Xi Jiang, Jack Luo +2
Low Latency, Low Loss, and Scalable Throughput (L4S), as an emerging router-queue management technique, has seen steady deployment in the industry. An L4S-enabled router assigns ea…
cs.CL2025
HyperRAG: Enhancing Quality-Efficiency Tradeoffs in Retrieval-Augmented Generation with Reranker KV-Cache Reuse
Yuwei An, Yihua Cheng, Seo Jin Park +1
Retrieval-Augmented Generation (RAG) has emerged as a powerful paradigm for enhancing the performance of large language models (LLMs) by integrating external knowledge into the gen…