1 paper
Wenxin Zhang, Yueying Li, Ciamac C. Moallemi +1
Prompt caching is critical for reducing latency and cost in LLM inference: OpenAI and Anthropic report up to 50-90% cost savings through prompt reuse. Despite its widespread succes…