Showing cs.DCShow all
2 papers · 1 filter
cs.DC2026
CacheWise: Understanding Workloads and Optimizing KVCache Management for Efficiently Serving LLM Coding Agents
Shubham Tiwari, Tapan Chugh, Nash Rickert +3
Coding agents are a fast-growing LLM application, executing as long-running closed-loop sessions in which LLM generations alternate with external tool calls. Yet, unlike chat workl…
cs.DC2025
Cloud abstractions for AI workloads
Marco Canini, Theophilus A. Benson, Ricardo Bianchini +4
AI workloads, often hosted in multi-tenant cloud environments, require vast computational resources but suffer inefficiencies due to limited tenant-provider coordination. Tenants l…