2 papers
cs.AR2026
Characterizing LLM Kernel Access and Memory Interaction in Multi-Partition NUMA GPUs
Donghyeon Joo, Sooraj Puthoor, Nuwan Jayasena +1
Large language model (LLM) workloads motivate multi-partition GPUs as a path to scaling compute and memory capacity, but their non-uniform memory access characteristics and inter-p…
cs.LG2025
Mustafar: Promoting Unstructured Sparsity for KV Cache Pruning in LLM Inference
Donghyeon Joo, Helya Hosseini, Ramyad Hadidi +1
We demonstrate that unstructured sparsity significantly improves KV cache compression for LLMs, enabling sparsity levels up to 70% without compromising accuracy or requiring fine-t…