16 citations · 28 across the 4 of their papers we have counts for
8 papers
Towards Load-Aware Prefill Deflection for Disaggregated LLM Serving
Shrikara Arun, Anjaly Parayil, Srikant Bharadwaj +2
Disaggregated LLM serving runs prefill and decode on separate GPU pools to keep the two phases from interfering. In practice, this creates a new asymmetry: under bursty, heavy-tail…
SIFT: Selective-Index For Fast Compute of RAG Prefill by Exploiting Attention Invariance
Rya Sanovar, Srikant Bharadwaj, Hritvik Taneja +1
Retrieval-Augmented Generation (RAG) injects LLM queries with relevant documents to improve response quality. This injection increases prompt length and slows time to first token (…
Taming the Tail: NoI Topology Synthesis for Mixed DL Workloads on Chiplet-Based Accelerators
Arnav Shukla, Harsh Sharma, Srikant Bharadwaj +2
Heterogeneous chiplet-based systems improve scaling by disag-gregating CPUs/GPUs and emerging technologies (HBM/DRAM).However this on-package disaggregation introduces a latency in…
TurboAttention: Efficient Attention Approximation For High Throughputs LLMs
Hao Kang, Srikant Bharadwaj, James Hensman +3
Large language model (LLM) inference demands significant amount of computation and memory, especially in the key attention mechanism. While techniques, such as quantization and acc…
Predict; Do not React for Enabling Efficient Fine Grain DVFS in GPUs
Srikant Bharadwaj, Shomit Das, Kaushik Mazumdar +2
With the continuous improvement of on-chip integrated voltage regulators (IVRs) and fast, adaptive frequency control, dynamic voltage-frequency scaling (DVFS) transition times have…
Accelerating Variational Quantum Algorithms Using Circuit Concurrency
Salonik Resch, Anthony Gutierrez, Joon Suk Huh +5
Variational quantum algorithms (VQAs) provide a promising approach to achieve quantum advantage in the noisy intermediate-scale quantum era. In this era, quantum computers experien…