Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
Short window attention enables long-term memorization
Loïc Cabannes, Maximilian Beck, Gergely Szilvasy +6
Recent works show that hybrid architectures combining local sliding window attention layers and global attention layers outperform either of these architectures taken separately. H…
cs.LG2025
Stochastic activations
Maria Lomeli, Matthijs Douze, Gergely Szilvasy +7
We introduce stochastic activations. This novel strategy randomly selects between several non-linear functions in the feed-forward layer of a large language model. In particular, w…