2 papers
cs.LG2024
Accumulator-Aware Post-Training Quantization for Large Language Models
Ian Colbert, Giuseppe Franco, Fabian Grob +2
When quantizing weights and activations to increasingly narrower representations, the cost of additions begins to dominate that of multiplications in multiply-accumulate (MAC) unit…
cs.IT2020
Faster Binary Embeddings for Preserving Euclidean Distances
Jinjie Zhang, Rayan Saab
We propose a fast, distance-preserving, binary embedding algorithm to transform a high-dimensional dataset into binary sequences in the cube $\{\…