1 paper · 1 filter
Yuze Liu, Yunhan Wang, Tiehua Zhang +5
The surge in intelligent applications driven by large language models (LLMs) has made it increasingly difficult for bandwidth-limited cloud servers to process extensive LLM workloa…