Showing cs.CLShow all
2 papers · 1 filter
cs.CL2026
Quant: Quantizing Language Models in Logarithmic Space
Jeremias Bohn, Tizian Dippold, Mahdi Koubaa +2
Quantization has become an invaluable tool to reduce memory requirements and inference speed of modern language models, in particular to make them available for consumer setups and…
cs.CL2025
The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
Alexander M. Fichtl, Jeremias Bohn, Josefin Kelber +2
Transformers have dominated sequence processing tasks for the past seven years -- most notably language modeling. However, the inherent quadratic complexity of their attention mech…