3 papers
cs.AR2025
MGS: Markov Greedy Sums for Accurate Low-Bitwidth Floating-Point Accumulation
Vikas Natesh, H. T. Kung, David Kong
We offer a novel approach, MGS (Markov Greedy Sums), to improve the accuracy of low-bitwidth floating-point dot products in neural network computations. In conventional 32-bit floa…
cs.LG2025
PQS (Prune, Quantize, and Sort): Low-Bitwidth Accumulation of Dot Products in Neural Network Computations
Vikas Natesh, H. T. Kung
We present PQS, which uses three techniques together - Prune, Quantize, and Sort - to achieve low-bitwidth accumulation of dot products in neural network computations. In conventio…
cs.LG2023
MEMA Runtime Framework: Minimizing External Memory Accesses for TinyML on Microcontrollers
Andrew Sabot, Vikas Natesh, H. T. Kung +1
We present the MEMA framework for the easy and quick derivation of efficient inference runtimes that minimize external memory accesses for matrix multiplication on TinyML systems.…