2 papers
cs.LG2025
Myosotis: structured computation for attention like layer
Evgenii Egorov, Hanno Ackermann, Markus Nagel +1
Attention layers apply a sequence-to-sequence mapping whose parameters depend on the pairwise interactions of the input elements. However, without any structural assumptions, memor…
cs.LG2025
GPTVQ: The Blessing of Dimensionality for LLM Quantization
Mart van Baalen, Andrey Kuzmin, Ivan Koryakovskiy +6
In this work we show that the size versus accuracy trade-off of neural network quantization can be significantly improved by increasing the quantization dimensionality. We propose…