2 papers
cs.LG2024
Attamba: Attending To Multi-Token States
Yash Akhauri, Safeen Huda, Mohamed S. Abdelfattah
When predicting the next token in a sequence, vanilla transformers compute attention over all previous tokens, resulting in quadratic scaling of compute with sequence length. State…
cs.LG2024
ShadowLLM: Predictor-based Contextual Sparsity for Large Language Models
Yash Akhauri, Ahmed F AbouElhamayed, Jordan Dotzel +4
The high power consumption and latency-sensitive deployments of large language models (LLMs) have motivated efficiency techniques like quantization and sparsity. Contextual sparsit…