6 citations · 7 across the 4 of their papers we have counts for
3 papers · 1 filter
Strata: Hierarchical Context Caching for Long Context Language Model Serving
Zhiqiang Xie, Ziyi Xu, Mark Zhao +5
Large Language Models (LLMs) with expanding context windows face significant performance hurdles. While caching key-value (KV) states is critical for avoiding redundant computation…
At-Scale Sparse Deep Neural Network Inference with Efficient GPU Implementation
Mert Hidayetoglu, Carl Pearson, Vikram Sharma Mailthody +4
This paper presents GPU performance optimization and scaling results for inference models of the Sparse Deep Neural Network Challenge 2020. Demands for network quality have increas…
EMOGI: Efficient Memory-access for Out-of-memory Graph-traversal In GPUs
Seung Won Min, Vikram Sharma Mailthody, Zaid Qureshi +3
Modern analytics and recommendation systems are increasingly based on graph data that capture the relations between entities being analyzed. Practical graphs come in huge sizes, of…