2 papers
cs.LG2025
ParaRNN: Unlocking Parallel Training of Nonlinear RNNs for Large Language Models
Federico Danieli, Pau Rodriguez, Miguel Sarabia +2
Recurrent Neural Networks (RNNs) laid the foundation for sequence modeling, but their intrinsic sequential nature restricts parallel computation, creating a fundamental barrier to…
cs.LG2025
Understanding Input Selectivity in Mamba: Impact on Approximation Power, Memorization, and Associative Recall Capacity
Ningyuan Huang, Miguel Sarabia, Abhinav Moudgil +3
State-Space Models (SSMs), and particularly Mamba, have recently emerged as a promising alternative to Transformers. Mamba introduces input selectivity to its SSM layer (S6) and in…