17 citations · 17 across the 2 of their papers we have counts for
4 papers · 1 filter
An Internet for the KV Cache: Rethinking Classical Infrastructure Boundaries in the LLM Inference Age
Siddhant Ray, Nick Feamster, Junchen Jiang
LLM inference has become a global-scale, heterogeneous workload spanning agents, retrieval, tool-use, code execution and multi-modal reasoning. These workloads naturally enable con…
SwiftQueue: Optimizing Low-Latency Applications with Swift Packet Queuing
Siddhant Ray, Xi Jiang, Jack Luo +2
Low Latency, Low Loss, and Scalable Throughput (L4S), as an emerging router-queue management technique, has seen steady deployment in the industry. An L4S-enabled router assigns ea…
NetLLM: Adapting Large Language Models for Networking
Duo Wu, Xianda Wang, Yaqi Qiao +4
Many networking tasks now employ deep learning (DL) to solve complex prediction and optimization problems. However, current design philosophy of DL-based algorithms entails intensi…
Eloquent: A More Robust Transmission Scheme for LLM Token Streaming
Hanchen Li, Yuhan Liu, Yihua Cheng +3
To render each generated token in real-time for users, the Large Language Model (LLM) server generates tokens one by one and streams each token (or group of a few tokens) through t…