10 papers
Price of metric universality in vector quantization is at most 0.11 bit
Alina Harbuzova, Or Ordentlich, Yury Polyanskiy
Fast computation of a matrix product is a workhorse of modern LLMs. To make their deployment more efficient, a popular approach is that of using a low-precision approxim…
High-Rate Quantized Matrix Multiplication II
Or Ordentlich, Yury Polyanskiy
This is the second part of the work investigating quantized matrix multiplication (MatMul). In part I we considered the case of calibration-free quantization, whereas here we discu…
WaterSIC: Information-Theoretically (Near) Optimal Linear Layer Quantization
Egor Lifar, Semyon Savkin, Or Ordentlich +1
This paper considers the problem of converting a given dense linear layer to low precision. The tradeoff between compressed length and output discrepancy is analyzed information th…
High-Rate Quantized Matrix Multiplication I
Or Ordentlich, Yury Polyanskiy
This paper investigates the problem of quantized matrix multiplication (MatMul), which has become crucial for the efficient deployment of large language models (LLMs). We consider…
The Voronoi Spherical CDF for Lattices and Linear Codes: New Bounds for Quantization and Coding
Or Ordentlich
For a lattice/linear code, we define the Voronoi spherical cumulative density function (CDF) as the CDF of the -norm/Hamming weight of a random vector uniformly distributed…
Optimal Quantization for Matrix Multiplication
Or Ordentlich, Yury Polyanskiy
Recent work in machine learning community proposed multiple methods for performing lossy compression (quantization) of large matrices. This quantization is important for accelerati…