14 citations · 38 across the 9 of their papers we have counts for
15 papers · 1 filter
Laughing Hyena Distillery: Extracting Compact Recurrences From Convolutions
Stefano Massaroli, Michael Poli, Daniel Y. Fu +11
Recent advances in attention-free sequence models rely on convolutions as alternatives to the attention operator at the core of Transformers. In particular, long convolution sequen…
Monarch Mixer: A Simple Sub-Quadratic GEMM-Based Architecture
Daniel Y. Fu, Simran Arora, Jessica Grogan +7
Machine learning models are increasingly being scaled in both sequence length and model dimension to reach longer contexts and better performance. However, existing architectures s…
HyenaDNA: Long-Range Genomic Sequence Modeling at Single Nucleotide Resolution
Eric Nguyen, Michael Poli, Marjan Faizi +10
Genomic (DNA) sequences encode an enormous amount of information for gene regulation and protein synthesis. Similar to natural language models, researchers have proposed foundation…
Transform Once: Efficient Operator Learning in Frequency Domain
Michael Poli, Stefano Massaroli, Federico Berto +4
Spectral analysis provides one of the most effective paradigms for information-preserving dimensionality reduction, as simple descriptions of naturally occurring signals are often…
Self-Similarity Priors: Neural Collages as Differentiable Fractal Representations
Michael Poli, Winnie Xu, Stefano Massaroli +3
Many patterns in nature exhibit self-similarity: they can be compactly described via self-referential transformations. Said patterns commonly appear in natural and artificial objec…
Monarch: Expressive Structured Matrices for Efficient and Accurate Training
Tri Dao, Beidi Chen, Nimit Sohoni +7
Large neural networks excel in many domains, but they are expensive to train and fine-tune. A popular approach to reduce their compute or memory requirements is to replace dense we…