4 citations · 4 across the 2 of their papers we have counts for
2 papers
cs.CL2025
Speculative Prefill: Turbocharging TTFT with Lightweight and Training-Free Token Importance Estimation
Jingyu Liu, Beidi Chen, Ce Zhang
Improving time-to-first-token (TTFT) is an essentially important objective in modern large language model (LLM) inference engines. Optimizing TTFT directly results in higher maxima…
cs.LG2023★ 4 cited
Laughing Hyena Distillery: Extracting Compact Recurrences From Convolutions
Stefano Massaroli, Michael Poli, Daniel Y. Fu +11
Recent advances in attention-free sequence models rely on convolutions as alternatives to the attention operator at the core of Transformers. In particular, long convolution sequen…