2 papers
cs.LG2026
The Curious Case of In-Training Compression of State Space Models
Makram Chahine, Philipp Nazari, Daniela Rus +1
State Space Models (SSMs), developed to tackle long sequence modeling tasks efficiently, offer both parallelizable training and fast inference. At their core are recurrent dynamica…
cs.LG2026
The Key to State Reduction in Linear Attention: A Rank-based Perspective
Philipp Nazari, T. Konstantin Rusch
Linear attention offers a computationally efficient yet expressive alternative to softmax attention. However, recent empirical results indicate that the hidden state of trained lin…