2 papers
cs.PF2025
Meta-Metrics and Best Practices for System-Level Inference Performance Benchmarking
Shweta Salaria, Zhuoran Liu, Nelson Mimura Gonzalez
Benchmarking inference performance (speed) of Foundation Models such as Large Language Models (LLM) involves navigating a vast experimental landscape to understand the complex inte…
cs.PF2025
Statistical Modeling and Uncertainty Estimation of LLM Inference Systems
Kaustabha Ray, Nelson Mimura Gonzalez, Bruno Wassermann +2
Large Language Model (LLM) inference systems present significant challenges in statistical performance characterization due to dynamic workload variations, diverse hardware archite…