1 paper · 1 filter
Quqing Zhang, Kai Chen, Ning Liao +5
In long-context LLM serving, the prefill stage often dominates time-to-first-token and computational cost. Although Prefix Cache in vLLM/PagedAttention has been widely used to reus…