6 papers
Characterizing LLM Kernel Access and Memory Interaction in Multi-Partition NUMA GPUs
Donghyeon Joo, Sooraj Puthoor, Nuwan Jayasena +1
Large language model (LLM) workloads motivate multi-partition GPUs as a path to scaling compute and memory capacity, but their non-uniform memory access characteristics and inter-p…
Tureis: Transformer-based Unified Resilience for IoT Devices in Smart Homes
Alireza Borhani, Vafa Andalibi, Bahar Asgari
Smart-home IoT systems rely on heterogeneous sensor networks whose correctness shapes application behavior and the physical environment. However, these low-cost, resource-constrain…
Arcalís: Accelerating Remote Procedure Calls Using a Líghtweight Near-Cache Solution
Johnson Umeike, Pongstorn Maidee, Bahar Asgari
Modern microservices increasingly depend on high-performance remote procedure calls (RPCs) to coordinate fine-grained, distributed computation. As network bandwidths continue to sc…
Mustafar: Promoting Unstructured Sparsity for KV Cache Pruning in LLM Inference
Donghyeon Joo, Helya Hosseini, Ramyad Hadidi +1
We demonstrate that unstructured sparsity significantly improves KV cache compression for LLMs, enabling sparsity levels up to 70% without compromising accuracy or requiring fine-t…
Belenos: Bottleneck Evaluation to Link Biomechanics to Novel Computing Optimizations
Hana Chitsaz, Johnson Umeike, Amirmahdi Namjoo +2
Finite element simulations are essential in biomechanics, enabling detailed modeling of tissues and organs. However, architectural inefficiencies in current hardware and software s…
GUST: Graph Edge-Coloring Utilization for Accelerating Sparse Matrix Vector Multiplication
Armin Gerami, Bahar Asgari
Sparse matrix-vector multiplication (SpMV) plays a vital role in various scientific and engineering fields, from scientific computing to machine learning. Traditional general-purpo…