3 citations · 3 across the 6 of their papers we have counts for
Showing cs.DCShow all
2 papers · 1 filter
cs.DC2026
RAPID-Serve: Resource-efficient and Accelerated P/D Intra-GPU Disaggregation
Amna Masood, Pratishtha Gaur, Nuwan Jayasena
Two widely adopted techniques for LLM inference serving systems today are hybrid batching and disaggregated serving. A hybrid batch combines prefill and decode tokens of different…
cs.DC2020
SeqPoint: Identifying Representative Iterations of Sequence-based Neural Networks
Suchita Pati, Shaizeen Aga, Matthew D. Sinclair +1
The ubiquity of deep neural networks (DNNs) continues to rise, making them a crucial application class for hardware optimizations. However, detailed profiling and characterization…