2 papers
cs.AI2026
SIFT: Selective-Index For Fast Compute of RAG Prefill by Exploiting Attention Invariance
Rya Sanovar, Srikant Bharadwaj, Hritvik Taneja +1
Retrieval-Augmented Generation (RAG) injects LLM queries with relevant documents to improve response quality. This injection increases prompt length and slows time to first token (…
cs.AR2025
Lean Attention: Hardware-Aware Scalable Attention Mechanism for the Decode-Phase of Transformers
Rya Sanovar, Srikant Bharadwaj, Renee St. Amant +2
Transformer-based models have emerged as one of the most widely used architectures for natural language processing, natural language generation, and image generation. The size of t…