Showing 2026Show all
2 papers · 1 filter
cs.CL2026
What Attention Recalls and Recurrence Controls in Hybrid Language Models
Kirill Afendulev, Alexey Dontsov, Elena Tutubalina +1
Hybrid language models combine attention with a fixed-size recurrent state, but the role of each channel remains unclear. We introduce two cache-level interventions. Split-prefill…
cs.LG2026
Sanity Checks for Sparse Autoencoders: Do SAEs Beat Random Baselines?
Anton Korznikov, Andrey Galichin, Alexey Dontsov +3
Sparse Autoencoders (SAEs) have emerged as a promising tool for interpreting neural networks by decomposing their activations into sparse sets of human-interpretable features. Rece…