14 citations · 24 across the 3 of their papers we have counts for
3 papers
cs.CL2024★ 1 cited
Just read twice: closing the recall gap for recurrent language models
Simran Arora, Aman Timalsina, Aaryan Singhal +6
Recurrent large language models that compete with Transformers in language modeling perplexity are emerging at a rapid rate (e.g., Mamba, RWKV). Excitingly, these architectures use…
cs.LG2023★ 14 cited
Monarch Mixer: A Simple Sub-Quadratic GEMM-Based Architecture
Daniel Y. Fu, Simran Arora, Jessica Grogan +7
Machine learning models are increasingly being scaled in both sequence length and model dimension to reach longer contexts and better performance. However, existing architectures s…
cs.AI2023★ 9 cited
Accelerating LLM Inference with Staged Speculative Decoding
Benjamin Spector, Chris Re
Recent advances with large language models (LLM) illustrate their diverse capabilities. We propose a novel algorithm, staged speculative decoding, to accelerate LLM inference in sm…