Showing cs.ARShow all
2 papers · 1 filter
cs.AR2026
Celty: SpMspV GPU Kernel and SIMT Co-Design for Efficient Dual-Sparse LLM Inference
Ruokai Yin, Priyadarshini Panda
Large Language Models (LLMs) increasingly rely on sparsity to reduce inference cost, but most prior work targets a single sparsity source-either weight or activation-and optimizes…
cs.AR2026
BRIM: Workload-Balanced Dual-Sided Bit-Serial Sparse Inference Accelerator
Varun Manjunath, Ruokai Yin, Donghyun Lee +2
Bit-serial accelerators exploit bit-level sparsity to reduce DNN inference cost, but existing designs exploit sparsity on only one operand, bounding the speedup. Extending sparsity…