47 citations · 47 across the 2 of their papers we have counts for
2 papers
cs.DC2026
RAPID-Serve: Resource-efficient and Accelerated P/D Intra-GPU Disaggregation
Amna Masood, Pratishtha Gaur, Nuwan Jayasena
Two widely adopted techniques for LLM inference serving systems today are hybrid batching and disaggregated serving. A hybrid batch combines prefill and decode tokens of different…
cs.DC2022★ 47 cited
GPU-Initiated On-Demand High-Throughput Storage Access in the BaM System Architecture
Zaid Qureshi, Vikram Sharma Mailthody, Isaac Gelado +10
Graphics Processing Units (GPUs) have traditionally relied on the host CPU to initiate access to the data storage. This approach is well-suited for GPU applications with known data…