2 papers
cs.DC2026
Revati: Transparent GPU-Free Time-Warp Emulation for LLM Serving
Amey Agrawal, Mayank Yadav, Sukrit Kumar +9
Deploying LLMs efficiently requires testing hundreds of serving configurations, but evaluating each one on a GPU cluster takes hours and costs thousands of dollars. Discrete-event…
cs.LG2025
Maya: Optimizing Deep Learning Training Workloads using GPU Runtime Emulation
Srihas Yarlagadda, Amey Agrawal, Elton Pinto +6
Training large foundation models costs hundreds of millions of dollars, making deployment optimization critical. Current approaches require machine learning engineers to manually c…