1 paper
Jing Ding, Yash Nishant, Chandrish Ambati +2
LLM inference at scale faces a memory wall. The KV cache demands tens of terabytes at hundreds of gigabytes per second, yet no current memory tier delivers both at once. Characteri…