3 papers
stat.ML2026
Throughput-Optimal Scheduling Algorithms for LLM Inference and AI Agents
J. G. Dai, Tianze Deng, Yueying Li +1
As demand for Large Language Models (LLMs) and AI agents grows rapidly, optimizing systems for efficient LLM inference becomes critical. While significant efforts have targeted sys…
eess.SY2025
Tail-Optimized Caching for LLM Inference
Wenxin Zhang, Yueying Li, Ciamac C. Moallemi +1
Prompt caching is critical for reducing latency and cost in LLM inference: OpenAI and Anthropic report up to 50-90% cost savings through prompt reuse. Despite its widespread succes…
cs.CL2025
SplitReason: Learning To Offload Reasoning
Yash Akhauri, Anthony Fei, Chi-Chih Chang +3
Reasoning in large language models (LLMs) tends to produce substantially longer token generation sequences than simpler language modeling tasks. This extended generation length ref…