3 citations · 4 across the 8 of their papers we have counts for
1 paper · 1 filter
Xuan He, Zequan Fang, Jinzhao Lian +3
The ever-increasing computation and energy demand for LLM and AI agents call for holistic and efficient optimization of LLM serving systems. In practice, heterogeneous GPU clusters…