57 citations · 57 across the 3 of their papers we have counts for
3 papers
cs.LG2025
SPIRe: Boosting LLM Inference Throughput with Speculative Decoding
Sanjit Neelam, Daniel Heinlein, Vaclav Cvicek +2
Speculative decoding (SD) has been shown to reduce the latency of autoregressive decoding (AD) by 2-3x for small batch sizes. However, increasing throughput and therefore reducing…
cs.CL2024
Naive Bayes-based Context Extension for Large Language Models
Jianlin Su, Murtadha Ahmed, Wenbo +3
Large Language Models (LLMs) have shown promising in-context learning abilities. However, conventional In-Context Learning (ICL) approaches are often impeded by length limitations…
cs.LG2022★ 57 cited
Efficiently Scaling Transformer Inference
Reiner Pope, Sholto Douglas, Aakanksha Chowdhery +7
We study the problem of efficient generative inference for Transformer models, in one of its most challenging settings: large deep models, with tight latency targets and long seque…