1 paper
Zixi Chen, Tianci Bu, Chendong Song +3
Over 40% of computational power in Large Language Model (LLM) serving systems can be systematically wasted - not from hardware limits, but from load imbalance in barrier-synchroniz…