Showing 2026Show all
2 papers · 1 filter
cs.LG2026
Provably Shorter Scratchpads in Hybrid DeltaNet-Attention Decoders
Tomasz Steifer
We investigate the expressive power of hybrid recurrent-attention decoders, a class of architectures used in recent open-source language models such as Qwen3-Next and its successor…
cs.LG2026
Parity, Sensitivity, and Transformers
Alexander Kozachinskiy, Tomasz Steifer, PrzemysÅaw WaÅÈ©ga
Understanding what neural architectures can and cannot compute is a central challenge in the theory of AI. One of the fundamental problems in this context is the PARITY task, which…