1 paper
Huamin Chen, Xunzhuo Liu, Yuhan Liu +3
How many tokens can a GPU inference cluster deliver per watt? Across deployments of identical hardware, the answer varies by 40x -- not because of software inefficiency, but becaus…