354 citations
- Microsoft (United States)US12 papers
- Indian Institute of Technology DelhiIN8 papers
- Institut national de recherche en sciences et technologies du numériqueFR8 papers
- Microsoft Research (United Kingdom)GB8 papers
- Indian Institute of Science BangaloreIN7 papers
- Carnegie Mellon UniversityUS5 papers
- Indian Institute of Technology KanpurIN5 papers
- Indian Institute of Technology KharagpurIN5 papers
- Birla Institute of Technology and Science, Pilani - Goa CampusIN4 papers
- University of Illinois Urbana-ChampaignUS4 papers
- University of WashingtonUS4 papers
- Cornell UniversityUS3 papers
Showing 2024 · cs.LGShow all
2 papers · 2 filters
cs.LG2024★ 20 cited
POD-Attention: Unlocking Full Prefill-Decode Overlap for Faster LLM Inference
Aditya K Kamath, Ramya Prabhu, Jayashree Mohan +3
Each request in LLM inference goes through two phases: compute-bound prefill and memory-bandwidth-bound decode. To improve GPU utilization, recent systems use hybrid batching that…
cs.LG2024★ 1 cited
Linear Contextual Bandits with Hybrid Payoff: Revisited
Nirjhar Das, Gaurav Sinha
We study the Linear Contextual Bandit problem in the hybrid reward setting. In this setting every arm's reward model contains arm specific parameters in addition to parameters shar…