papers

Publications (50)

cs.LG2025

On the Benefits of Memory for Modeling Time-Dependent PDEs

Ricardo Buitrago Ruiz, Tanya Marwah, Albert Gu +1

cs.LG2022

Diagonal State Spaces are as Effective as Structured State Spaces

Ankit Gupta, Albert Gu, Jonathan Berant

cs.CL2023

Augmenting conformers with structured state-space sequence models for online speech recognition

Haozhe Shan, Albert Gu, Zhong Meng +3

cs.LG2024

Hydra: Bidirectional State Space Models Through Generalized Matrix Mixers

Sukjun Hwang, Aakash Lahoti, Tri Dao +1

cs.LG2023

Resurrecting Recurrent Neural Networks for Long Sequences

Antonio Orvieto, Samuel L Smith, Albert Gu +4

cs.LG2026

dnaHNet: A Scalable and Hierarchical Foundation Model for Genomic Sequence Learning

Arnav Shah, Junzhe Li, Parsa Idehpour +7

cs.CV2022

S4ND: Modeling Images and Videos as Multidimensional Signals Using State Spaces

Eric Nguyen, Karan Goel, Albert Gu +5

cs.DS2013

The Power of Deferral: Maintaining a Constant-Competitive Steiner Tree Online

Albert Gu, Anupam Gupta, Amit Kumar

cs.CV2025

Autoregressive Universal Video Segmentation Model

Miran Heo, Sukjun Hwang, Min-Hung Chen +4

cs.LG2025

Understanding and Improving Length Generalization in Recurrent Models

Ricardo Buitrago Ruiz, Albert Gu

cs.CL2025

Thinking Slow, Fast: Scaling Inference Compute with Distilled Reasoners

Daniele Paliotta, Junxiong Wang, Matteo Pagliardini +6

cs.LG2025

Dynamic Chunking for End-to-End Hierarchical Sequence Modeling

Sukjun Hwang, Brandon Wang, Albert Gu

cs.LG2024

Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Duality

Tri Dao, Albert Gu

cs.LG2026

Mamba-3: Improved Sequence Modeling using State Space Principles

Aakash Lahoti, Kevin Y. Li, Berlin Chen +5

cs.CV2023

Modelling Long Range Dependencies in D: From Task-Specific to a General Purpose CNN

David M. Knigge, David W. Romero, Albert Gu +5

cs.NE2020

Improving the Gating Mechanism of Recurrent Neural Networks

Albert Gu, Caglar Gulcehre, Tom Le Paine +2

cs.DS2020

From Trees to Continuous Embeddings and Back: Hyperbolic Hierarchical Clustering

Ines Chami, Albert Gu, Vaggos Chatziafratis +1

cs.LG2021

Combining Recurrent, Convolutional, and Continuous-time Models with Linear State-Space Layers

Albert Gu, Isys Johnson, Karan Goel +4

cs.LG2021

HoroPCA: Hyperbolic Dimensionality Reduction via Horospherical Projections

Ines Chami, Albert Gu, Dat Nguyen +1

cs.CL2023

Pretraining Without Attention

Junxiong Wang, Jing Nathan Yan, Albert Gu +1

cs.LG2025

Llamba: Scaling Distilled Recurrent Models for Efficient Language Processing

Aviv Bick, Tobias Katsch, Nimit Sohoni +2

cs.LG2022

Towards a General Purpose CNN for Long Range Dependencies in D

David W. Romero, David M. Knigge, Albert Gu +4

cs.SD2022

It's Raw! Audio Generation with State-Space Models

Karan Goel, Albert Gu, Chris Donahue +1

cs.LG2020

HiPPO: Recurrent Memory with Optimal Polynomial Projections

Albert Gu, Tri Dao, Stefano Ermon +2

cs.LG2024

An Empirical Study of Mamba-based Language Models

Roger Waleffe, Wonmin Byeon, Duncan Riach +13

cs.LG2025

Chimera: State Space Models Beyond Sequences

Aakash Lahoti, Tanya Marwah, Ratish Puduppully +1

cs.LG2026

Raven: High-Recall Sequence Modeling with Sparse Memory Routing

Arshia Afzal, Aviv Bick, Eric P. Xing +2

Raven is a linear-time sequence model that uses learned, input-dependent routing to update only a subset of fixed memory slots, reducing interference and improving long-range recal…

#sequence modeling#long-context recall#sparse memory#linear-time models
cs.LG2025

Understanding the Skill Gap in Recurrent Language Models: The Role of the Gather-and-Aggregate Mechanism

Aviv Bick, Eric Xing, Albert Gu

cs.LG2024

Mamba: Linear-Time Sequence Modeling with Selective State Spaces

Albert Gu, Tri Dao

cs.LG2023

Structured State Space Models for In-Context Reinforcement Learning

Chris Lu, Yannick Schroecker, Albert Gu +4

cs.LG2025

ARC-AGI Without Pretraining

Isaac Liao, Albert Gu

cs.LG2022

How to Train Your HiPPO: State Space Models with Generalized Orthogonal Basis Projections

Albert Gu, Isys Johnson, Aman Timalsina +2

cs.LG2024

Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models

Soham De, Samuel L. Smith, Anushan Fernando +14

cs.LG2022

No Subclass Left Behind: Fine-Grained Robustness in Coarse-Grained Classification Problems

Nimit S. Sohoni, Jared A. Dunnmon, Geoffrey Angus +2

cs.LG2019

Learning Compressed Transforms with Low Displacement Rank

Anna T. Thomas, Albert Gu, Tri Dao +2

cs.DS2017

A Two Pronged Progress in Structured Dense Matrix Multiplication

Christopher De Sa, Albert Gu, Rohan Puttagunta +2

cs.LG2026

Joint Distillation for Fast Likelihood Evaluation and Sampling in Flow-based Models

Xinyue Ai, Yutong He, Albert Gu +4

cs.LG2025

HybriDNA: A Hybrid Transformer-Mamba2 Long-Range DNA Language Model

Mingqian Ma, Guoqing Liu, Chuan Cao +12

cs.DS2019

Sparse Recovery for Orthogonal Polynomial Transforms

Anna Gilbert, Albert Gu, Christopher Re +2

cs.LG2019

A Kernel Theory of Modern Data Augmentation

Tri Dao, Albert Gu, Alexander J. Ratner +3

cs.LG2020

Learning Fast Algorithms for Linear Transforms Using Butterfly Factorizations

Tri Dao, Albert Gu, Matthew Eichhorn +2

cs.LG2026

Retrieval-Aware Distillation for Transformer-SSM Hybrids

Aviv Bick, Eric P. Xing, Albert Gu

cs.LG2022

Efficiently Modeling Long Sequences with Structured State Spaces

Albert Gu, Karan Goel, Christopher Ré

q-bio.GN2024

Caduceus: Bi-Directional Equivariant Long-Range DNA Sequence Modeling

Yair Schiff, Chia-Hsiang Kao, Aaron Gokaslan +3

cs.LG2021

Kaleidoscope: An Efficient, Learnable Representation For All Structured Linear Maps

Tri Dao, Nimit S. Sohoni, Albert Gu +5

cs.LG2025

Lyra: An Efficient and Expressive Subquadratic Architecture for Modeling Biological Sequences

Krithik Ramesh, Sameed M. Siddiqui, Albert Gu +2

cs.LG2025

Transformers to SSMs: Distilling Quadratic Knowledge to Subquadratic Models

Aviv Bick, Kevin Y. Li, Eric P. Xing +2

cs.LG2022

On the Parameterization and Initialization of Diagonal State Space Models

Albert Gu, Ankit Gupta, Karan Goel +1

cs.LG2020

Model Patching: Closing the Subgroup Performance Gap with Data Augmentation

Karan Goel, Albert Gu, Yixuan Li +1

cs.LG2018

Representation Tradeoffs for Hyperbolic Embeddings

Christopher De Sa, Albert Gu, Christopher Ré +1