1 paper
Longshaokan Wang, Wai Tsang Keung, Punit Ghodasara +3
Scaling LLM-based applications to millions of users is bottlenecked by the inference cost and latency of modern foundation models. A natural fix is to cluster the inputs and call t…