4 citations · 7 across the 4 of their papers we have counts for
4 papers
CHAI: Clustered Head Attention for Efficient LLM Inference
Saurabh Agarwal, Bilge Acun, Basil Hosmer +5
Large Language Models (LLMs) with hundreds of billions of parameters have transformed the field of machine learning. However, serving these models at inference time is both compute…
Not All GPUs Are Created Equal: Characterizing Variability in Large-Scale, Accelerator-Rich Systems
Prasoon Sinha, Akhil Guliani, Rutwik Jain +3
Scientists are increasingly exploring and utilizing the massive parallelism of general-purpose accelerators such as GPUs for scientific breakthroughs. As a result, datacenters, hyp…
KeystoneML: Optimizing Pipelines for Large-Scale Advanced Analytics
Evan R. Sparks, Shivaram Venkataraman, Tomer Kaftan +2
Modern advanced analytics applications make use of machine learning techniques and contain multiple steps of domain-specific and general-purpose processing with high resource requi…
Probabilistically Bounded Staleness for Practical Partial Quorums
Peter Bailis, Shivaram Venkataraman, Michael J. Franklin +2
Data store replication results in a fundamental trade-off between operation latency and data consistency. In this paper, we examine this trade-off in the context of quorum-replicat…