1 paper
Longshaokan Wang, Wai Tsang Keung, Punit Ghodasara +3
LLMs can be prohibitively expensive and slow to run at scale, especially for applications that invoke an LLM per sample over millions of inputs. A natural way to scale is to cluste…