45 citations · 142 across the 34 of their papers we have counts for
Showing 2024 · cs.LGShow all
2 papers · 2 filters
cs.LG2024★ 2 cited
Loki: Low-rank Keys for Efficient Sparse Attention
Prajwal Singhania, Siddharth Singh, Shwai He +2
Inference on large language models (LLMs) can be expensive in terms of the compute and memory costs involved, especially when long sequence lengths are used. In particular, the sel…
cs.LG2024★ 6 cited
Transformers Can Do Arithmetic with the Right Embeddings
Sean McLeish, Arpit Bansal, Alex Stein +8
The poor performance of transformers on arithmetic tasks seems to stem in large part from their inability to keep track of the exact position of each digit inside of a large span o…