2 papers
cs.AR2025
MGS: Markov Greedy Sums for Accurate Low-Bitwidth Floating-Point Accumulation
Vikas Natesh, H. T. Kung, David Kong
We offer a novel approach, MGS (Markov Greedy Sums), to improve the accuracy of low-bitwidth floating-point dot products in neural network computations. In conventional 32-bit floa…
cs.LG2025
PQS (Prune, Quantize, and Sort): Low-Bitwidth Accumulation of Dot Products in Neural Network Computations
Vikas Natesh, H. T. Kung
We present PQS, which uses three techniques together - Prune, Quantize, and Sort - to achieve low-bitwidth accumulation of dot products in neural network computations. In conventio…