1 citations · 1 across the 5 of their papers we have counts for
1 paper · 1 filter
Foteini Strati, Sara Mcallister, Amar Phanishayee +2
Distributed LLM serving is costly and often underutilizes hardware accelerators due to three key challenges: bubbles in pipeline-parallel deployments caused by the bimodal latency…