1 citations · 1 across the 2 of their papers we have counts for
8 papers
LayoutBench: Performance Benchmarking of Cloud Storage Layouts for Multimedia Data
Debopam Sanyal, Hongjie Chen, Alexey Tumanov +1
Modern multimedia machine learning workloads increasingly store large-scale datasets in cloud object storage services such as AWS S3. How these samples are physically organized in…
Revati: Transparent GPU-Free Time-Warp Emulation for LLM Serving
Amey Agrawal, Mayank Yadav, Sukrit Kumar +9
Deploying LLMs efficiently requires testing hundreds of serving configurations, but evaluating each one on a GPU cluster takes hours and costs thousands of dollars. Discrete-event…
No Request Left Behind: Tackling Heterogeneity in Long-Context LLM Inference with Medha
Amey Agrawal, Haoran Qiu, Junda Chen +6
Deploying million-token Large Language Models (LLMs) is challenging because production workloads are highly heterogeneous, mixing short queries and long documents. This heterogenei…
Maya: Optimizing Deep Learning Training Workloads using GPU Runtime Emulation
Srihas Yarlagadda, Amey Agrawal, Elton Pinto +6
Training large foundation models costs hundreds of millions of dollars, making deployment optimization critical. Current approaches require machine learning engineers to manually c…
On Evaluating Performance of LLM Inference Serving Systems
Amey Agrawal, Nitin Kedia, Anmol Agarwal +5
The rapid evolution of Large Language Model (LLM) inference systems has yielded significant efficiency improvements. However, our systematic analysis reveals that current evaluatio…
Etalon: Holistic Performance Evaluation Framework for LLM Inference Systems
Amey Agrawal, Anmol Agarwal, Nitin Kedia +5
Serving large language models (LLMs) in production can incur substantial costs, which has prompted recent advances in inference system optimizations. Today, these systems are evalu…