Showing cs.LGShow all
3 papers · 1 filter
cs.LG2025
PQS (Prune, Quantize, and Sort): Low-Bitwidth Accumulation of Dot Products in Neural Network Computations
Vikas Natesh, H. T. Kung
We present PQS, which uses three techniques together - Prune, Quantize, and Sort - to achieve low-bitwidth accumulation of dot products in neural network computations. In conventio…
cs.LG2023
Rosko: Row Skipping Outer Products for Sparse Matrix Multiplication Kernels
Vikas Natesh, Andrew Sabot, H. T. Kung +1
We propose Rosko -- row skipping outer products -- for deriving sparse matrix multiplication (SpMM) kernels in reducing computation and memory access requirements of deep neural ne…
cs.LG2023
MEMA Runtime Framework: Minimizing External Memory Accesses for TinyML on Microcontrollers
Andrew Sabot, Vikas Natesh, H. T. Kung +1
We present the MEMA framework for the easy and quick derivation of efficient inference runtimes that minimize external memory accesses for matrix multiplication on TinyML systems.…