10 papers · 1 filter
High-Rate Quantized Matrix Multiplication II
Or Ordentlich, Yury Polyanskiy
This is the second part of the work investigating quantized matrix multiplication (MatMul). In part I we considered the case of calibration-free quantization, whereas here we discu…
WaterSIC: Information-Theoretically (Near) Optimal Linear Layer Quantization
Egor Lifar, Semyon Savkin, Or Ordentlich +1
This paper considers the problem of converting a given dense linear layer to low precision. The tradeoff between compressed length and output discrepancy is analyzed information th…
Measure-to-measure Regression with Transformers
Matthew Vandergrift, Martha White, Yury Polyanskiy +2
Many learning problems require predicting how populations evolve under an unknown transformation. A natural representation for such populations is a probability measure, with point…
Representation Alignment Rests on Linear Structure
Kiril Bangachev, Guy Bresler, Yury Polyanskiy
We investigate the Platonic Representation Hypothesis (PRH) through a tripartite statistical framework of representations: signal, bias, and noise. {1) Signal:} We propose that Pla…
Is Dimensionality a Barrier for Retrieval Models?
Kiril Bangachev, Guy Bresler, Jonathan Kogan +1
Why does the low dimensionality of representations, typically , not prevent modern embedding-based retrieval models from scaling to billions, or even trillions, of d…
Continuous First, Discrete Later: VQ-VAEs Without Dimensional Collapse
Xinyu Zhao, Nikita Karagodin, Hamed Hassani +3
While many approaches to improve VQ-VAE performance focus on codebook size and utilization, the effect of dimensional collapse, where trained VQ-VAE representations live in an extr…