14 citations · 20 across the 3 of their papers we have counts for
3 papers
STAR: Synthesis of Tailored Architectures
Armin W. Thomas, Rom Parnichkun, Alexander Amini +2
Iterative improvement of model architectures is fundamental to deep learning: Transformers first enabled scaling, and recent advances in model hybridization have pushed the quality…
Monarch Mixer: A Simple Sub-Quadratic GEMM-Based Architecture
Daniel Y. Fu, Simran Arora, Jessica Grogan +7
Machine learning models are increasingly being scaled in both sequence length and model dimension to reach longer contexts and better performance. However, existing architectures s…
Simple Hardware-Efficient Long Convolutions for Sequence Modeling
Daniel Y. Fu, Elliot L. Epstein, Eric Nguyen +5
State space models (SSMs) have high performance on long sequence modeling but require sophisticated initialization techniques and specialized implementations for high quality and r…