6 citations · 6 across the 2 of their papers we have counts for
2 papers
cs.CL2024
PipeInfer: Accelerating LLM Inference using Asynchronous Pipelined Speculation
Branden Butler, Sixing Yu, Arya Mazaheri +1
Inference of Large Language Models (LLMs) across computer clusters has become a focal point of research in recent times, with many acceleration techniques taking inspiration from C…
cs.LG2024★ 6 cited
The Landscape and Challenges of HPC Research and LLMs
Le Chen, Nesreen K. Ahmed, Akash Dutta +14
Recently, language models (LMs), especially large language models (LLMs), have revolutionized the field of deep learning. Both encoder-decoder models and prompt-based techniques ha…