3 papers
cs.CL2026
Quant: Quantizing Language Models in Logarithmic Space
Jeremias Bohn, Tizian Dippold, Mahdi Koubaa +2
Quantization has become an invaluable tool to reduce memory requirements and inference speed of modern language models, in particular to make them available for consumer setups and…
cs.AI2026
Depth-Recurrent Attention Mixtures: Giving Latent Reasoning the Attention it Deserves
Jonas Knupp, Jan Hendrik Metzen, Jeremias Bohn +2
Depth-recurrence facilitates latent reasoning by sharing parameters across depths. However, prior work lacks combined FLOP-, parameter-, and memory-matched baselines, underutilizes…
cs.CL2025
The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
Alexander M. Fichtl, Jeremias Bohn, Josefin Kelber +2
Transformers have dominated sequence processing tasks for the past seven years -- most notably language modeling. However, the inherent quadratic complexity of their attention mech…