1 paper
Donghyeon Joo, Sooraj Puthoor, Nuwan Jayasena +1
Large language model (LLM) workloads motivate multi-partition GPUs as a path to scaling compute and memory capacity, but their non-uniform memory access characteristics and inter-p…