1 paper · 1 filter
Vimal William, Ravi Tandon, Jyotikrishna Dass
As Large Language Models scale to increasingly long contexts, the memory I/O and computational overhead of the Key-Value (KV) cache during decoding emerges as the primary throughpu…