2 papers
cs.LG2024
Approximate Top- for Increased Parallelism
Oscar Key, Luka Ribar, Alberto Cattaneo +2
We present an evaluation of bucketed approximate top- algorithms. Computing top- exactly suffers from limited parallelism, because the largest values must be aggregated a…
cs.LG2024
SparQ Attention: Bandwidth-Efficient LLM Inference
Luka Ribar, Ivan Chelombiev, Luke Hudlass-Galley +3
The computational difficulties of large language model (LLM) inference remain a significant obstacle to their widespread deployment. The need for many applications to support long…