Showing cs.ARShow all
3 papers · 1 filter
cs.AR2026
Characterizing LLM Kernel Access and Memory Interaction in Multi-Partition NUMA GPUs
Donghyeon Joo, Sooraj Puthoor, Nuwan Jayasena +1
Large language model (LLM) workloads motivate multi-partition GPUs as a path to scaling compute and memory capacity, but their non-uniform memory access characteristics and inter-p…
cs.AR2024
Global Optimizations & Lightweight Dynamic Logic for Concurrency
Suchita Pati, Shaizeen Aga, Nuwan Jayasena +1
Modern accelerators like GPUs are increasingly executing independent operations concurrently to improve the device's compute utilization. However, effectively harnessing it on GPUs…
cs.AR2024
T3: Transparent Tracking & Triggering for Fine-grained Overlap of Compute & Collectives
Suchita Pati, Shaizeen Aga, Mahzabeen Islam +2
Large Language Models increasingly rely on distributed techniques for their training and inference. These techniques require communication across devices which can reduce scaling e…