1 citations · 1 across the 2 of their papers we have counts for
1 paper · 1 filter
Xingqi Cui, Chieh-Jan Mike Liang, Jiarong Xing +1
Serving large generative models such as LLMs and multi- modal transformers requires balancing user-facing SLOs (e.g., time-to-first-token, time-between-tokens) with provider goals…