3 papers
cs.CL2026
Attention to Mamba: A Recipe for Cross-Architecture Distillation
Abhinav Moudgil, Ningyuan Huang, Eeshan Gunesh Dhekane +3
State Space Models (SSMs) such as Mamba have become a popular alternative to Transformer models, due to their reduced memory consumption and higher throughput at generation compare…
math.NA2025
Space-Time Block Preconditioning for Incompressible Resistive Magnetohydrodynamics
Federico Danieli, Ben S. Southworth, Jacob B. Schroder
This work develops an all-at-once space-time preconditioning approach for resistive magnetohydrodynamics (MHD). We consider parallel-in-time due to the long time domains required t…
cs.LG2025
Theory, Analysis, and Best Practices for Sigmoid Self-Attention
Jason Ramapuram, Federico Danieli, Eeshan Dhekane +8
Attention is a key part of the transformer architecture. It is a sequence-to-sequence mapping that transforms each sequence element into a weighted sum of values. The weights are t…