3 papers
cs.AR2026
Characterizing LLM Kernel Access and Memory Interaction in Multi-Partition NUMA GPUs
Donghyeon Joo, Sooraj Puthoor, Nuwan Jayasena +1
Large language model (LLM) workloads motivate multi-partition GPUs as a path to scaling compute and memory capacity, but their non-uniform memory access characteristics and inter-p…
cs.LG2025
Mustafar: Promoting Unstructured Sparsity for KV Cache Pruning in LLM Inference
Donghyeon Joo, Helya Hosseini, Ramyad Hadidi +1
We demonstrate that unstructured sparsity significantly improves KV cache compression for LLMs, enabling sparsity levels up to 70% without compromising accuracy or requiring fine-t…
cs.CL2024
Endor: Hardware-Friendly Sparse Format for Offloaded LLM Inference
Donghyeon Joo, Ramyad Hadidi, Soheil Feizi +1
The increasing size of large language models (LLMs) challenges their usage on resource-constrained platforms. For example, memory on modern GPUs is insufficient to hold LLMs that a…