1 paper
Joseph Kanichai, Tiziano De Matteis, Animesh Trivedi
Prefix caching can reduce the time to first token (TTFT) of long-context LLM requests by reusing previously computed key-value (KV) states, but for short prefixes or fast GPUs, rec…