Publications (50)
On the Benefits of Memory for Modeling Time-Dependent PDEs
Ricardo Buitrago Ruiz, Tanya Marwah, Albert Gu +1
Diagonal State Spaces are as Effective as Structured State Spaces
Ankit Gupta, Albert Gu, Jonathan Berant
Augmenting conformers with structured state-space sequence models for online speech recognition
Haozhe Shan, Albert Gu, Zhong Meng +3
Hydra: Bidirectional State Space Models Through Generalized Matrix Mixers
Sukjun Hwang, Aakash Lahoti, Tri Dao +1
Resurrecting Recurrent Neural Networks for Long Sequences
Antonio Orvieto, Samuel L Smith, Albert Gu +4
dnaHNet: A Scalable and Hierarchical Foundation Model for Genomic Sequence Learning
Arnav Shah, Junzhe Li, Parsa Idehpour +7
S4ND: Modeling Images and Videos as Multidimensional Signals Using State Spaces
Eric Nguyen, Karan Goel, Albert Gu +5
The Power of Deferral: Maintaining a Constant-Competitive Steiner Tree Online
Albert Gu, Anupam Gupta, Amit Kumar
Autoregressive Universal Video Segmentation Model
Miran Heo, Sukjun Hwang, Min-Hung Chen +4
Understanding and Improving Length Generalization in Recurrent Models
Ricardo Buitrago Ruiz, Albert Gu
Thinking Slow, Fast: Scaling Inference Compute with Distilled Reasoners
Daniele Paliotta, Junxiong Wang, Matteo Pagliardini +6
Dynamic Chunking for End-to-End Hierarchical Sequence Modeling
Sukjun Hwang, Brandon Wang, Albert Gu
Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Duality
Tri Dao, Albert Gu
Mamba-3: Improved Sequence Modeling using State Space Principles
Aakash Lahoti, Kevin Y. Li, Berlin Chen +5
Modelling Long Range Dependencies in D: From Task-Specific to a General Purpose CNN
David M. Knigge, David W. Romero, Albert Gu +5
Improving the Gating Mechanism of Recurrent Neural Networks
Albert Gu, Caglar Gulcehre, Tom Le Paine +2
From Trees to Continuous Embeddings and Back: Hyperbolic Hierarchical Clustering
Ines Chami, Albert Gu, Vaggos Chatziafratis +1
Combining Recurrent, Convolutional, and Continuous-time Models with Linear State-Space Layers
Albert Gu, Isys Johnson, Karan Goel +4
HoroPCA: Hyperbolic Dimensionality Reduction via Horospherical Projections
Ines Chami, Albert Gu, Dat Nguyen +1
Pretraining Without Attention
Junxiong Wang, Jing Nathan Yan, Albert Gu +1
Llamba: Scaling Distilled Recurrent Models for Efficient Language Processing
Aviv Bick, Tobias Katsch, Nimit Sohoni +2
Towards a General Purpose CNN for Long Range Dependencies in D
David W. Romero, David M. Knigge, Albert Gu +4
It's Raw! Audio Generation with State-Space Models
Karan Goel, Albert Gu, Chris Donahue +1
HiPPO: Recurrent Memory with Optimal Polynomial Projections
Albert Gu, Tri Dao, Stefano Ermon +2
An Empirical Study of Mamba-based Language Models
Roger Waleffe, Wonmin Byeon, Duncan Riach +13
Chimera: State Space Models Beyond Sequences
Aakash Lahoti, Tanya Marwah, Ratish Puduppully +1
Raven: High-Recall Sequence Modeling with Sparse Memory Routing
Arshia Afzal, Aviv Bick, Eric P. Xing +2
Raven is a linear-time sequence model that uses learned, input-dependent routing to update only a subset of fixed memory slots, reducing interference and improving long-range recal…
Understanding the Skill Gap in Recurrent Language Models: The Role of the Gather-and-Aggregate Mechanism
Aviv Bick, Eric Xing, Albert Gu
Mamba: Linear-Time Sequence Modeling with Selective State Spaces
Albert Gu, Tri Dao
Structured State Space Models for In-Context Reinforcement Learning
Chris Lu, Yannick Schroecker, Albert Gu +4
ARC-AGI Without Pretraining
Isaac Liao, Albert Gu
How to Train Your HiPPO: State Space Models with Generalized Orthogonal Basis Projections
Albert Gu, Isys Johnson, Aman Timalsina +2
Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models
Soham De, Samuel L. Smith, Anushan Fernando +14
No Subclass Left Behind: Fine-Grained Robustness in Coarse-Grained Classification Problems
Nimit S. Sohoni, Jared A. Dunnmon, Geoffrey Angus +2
Learning Compressed Transforms with Low Displacement Rank
Anna T. Thomas, Albert Gu, Tri Dao +2
A Two Pronged Progress in Structured Dense Matrix Multiplication
Christopher De Sa, Albert Gu, Rohan Puttagunta +2
Joint Distillation for Fast Likelihood Evaluation and Sampling in Flow-based Models
Xinyue Ai, Yutong He, Albert Gu +4
HybriDNA: A Hybrid Transformer-Mamba2 Long-Range DNA Language Model
Mingqian Ma, Guoqing Liu, Chuan Cao +12
Sparse Recovery for Orthogonal Polynomial Transforms
Anna Gilbert, Albert Gu, Christopher Re +2
A Kernel Theory of Modern Data Augmentation
Tri Dao, Albert Gu, Alexander J. Ratner +3
Learning Fast Algorithms for Linear Transforms Using Butterfly Factorizations
Tri Dao, Albert Gu, Matthew Eichhorn +2
Retrieval-Aware Distillation for Transformer-SSM Hybrids
Aviv Bick, Eric P. Xing, Albert Gu
Efficiently Modeling Long Sequences with Structured State Spaces
Albert Gu, Karan Goel, Christopher Ré
Caduceus: Bi-Directional Equivariant Long-Range DNA Sequence Modeling
Yair Schiff, Chia-Hsiang Kao, Aaron Gokaslan +3
Kaleidoscope: An Efficient, Learnable Representation For All Structured Linear Maps
Tri Dao, Nimit S. Sohoni, Albert Gu +5
Lyra: An Efficient and Expressive Subquadratic Architecture for Modeling Biological Sequences
Krithik Ramesh, Sameed M. Siddiqui, Albert Gu +2
Transformers to SSMs: Distilling Quadratic Knowledge to Subquadratic Models
Aviv Bick, Kevin Y. Li, Eric P. Xing +2
On the Parameterization and Initialization of Diagonal State Space Models
Albert Gu, Ankit Gupta, Karan Goel +1
Model Patching: Closing the Subgroup Performance Gap with Data Augmentation
Karan Goel, Albert Gu, Yixuan Li +1
Representation Tradeoffs for Hyperbolic Embeddings
Christopher De Sa, Albert Gu, Christopher Ré +1