1 paper
Geraldo F. Oliveira, Arash Tavakkol, Xiangyu Zhu +10
LLM inference is increasingly constrained by accelerator memory capacity rather than compute throughput. This constraint is especially acute in single-accelerator and small-node in…