1 paper · 1 filter
Srikanta Datta Tumkur, Mehar Simhadri, Anshu Bansal +5
When an LLM serving deployment runs out of KVcache room, there are two well-established ways out. Tensor parallelism shards the weights and the KV cache across two, four, or eight…