2 papers
cs.LG2025
Polar Sparsity: High Throughput Batched LLM Inferencing with Scalable Contextual Sparsity
Susav Shrestha, Brad Settlemyer, Nikoli Dryden +1
Accelerating large language model (LLM) inference is critical for real-world deployments requiring high throughput and low latency. Contextual sparsity, where each token dynamicall…
cs.IR2023
ESPN: Memory-Efficient Multi-Vector Information Retrieval
Susav Shrestha, Narasimha Reddy, Zongwang Li
Recent advances in large language models have demonstrated remarkable effectiveness in information retrieval (IR) tasks. While many neural IR systems encode queries and documents i…