1 paper · 1 filter
Wenzong Yang, Danyang Zhang, Kun Cao +17
The popularity of large language models (LLMs) escalates an ongoing demand for effective inference. However, due to the sequential processing of tokens during the token phase in de…