most citedLaughing Hyena Distillery: Extracting Compact Recurrences From Convolutions

4 citations · 5 across the 2 of their papers we have counts for

collaborators

5 papers

cs.CL20241 cited

Just read twice: closing the recall gap for recurrent language models

Simran Arora, Aman Timalsina, Aaryan Singhal +6

Recurrent large language models that compete with Transformers in language modeling perplexity are emerging at a rapid rate (e.g., Mamba, RWKV). Excitingly, these architectures use…

math.AT2024

Computing Generalized Ranks of Persistence Modules via Unfolding to Zigzag Modules

Tamal K. Dey, Cheng Xin

For a -indexed persistence module , the (generalized) rank of is defined as the rank of the limit-to-colimit map for the diagram of vector spaces of

cs.CL2024

Simple linear attention language models balance the recall-throughput tradeoff

Simran Arora, Sabri Eyuboglu, Michael Zhang +6

Recent work has shown that attention-based language models excel at recall, the ability to ground generations in tokens previously seen in context. However, the efficiency of atten…

cs.CL2023

Zoology: Measuring and Improving Recall in Efficient Language Models

Simran Arora, Sabri Eyuboglu, Aman Timalsina +5

Attention-free language models that combine gating and convolutions are growing in popularity due to their efficiency and increasingly competitive performance. To better understand…

cs.LG20234 cited

Laughing Hyena Distillery: Extracting Compact Recurrences From Convolutions

Stefano Massaroli, Michael Poli, Daniel Y. Fu +11

Recent advances in attention-free sequence models rely on convolutions as alternatives to the attention operator at the core of Transformers. In particular, long convolution sequen…