2 citations · 4 across the 11 of their papers we have counts for
1 paper · 1 filter
Yihao Zhao, Jiadun Chen, Peng Sun +3
Large language models (LLMs) with different architectures and sizes have been developed. Serving each LLM with dedicated GPUs leads to resource waste and service inefficiency due t…