1 paper
Matthew Adiletta, Gu-Yeon Wei, David Brooks
Large language model (LLM) inference performance is increasingly bottlenecked by the memory wall. While GPUs continue to scale raw compute throughput, they struggle to deliver scal…