1 paper
Haoyu Chen, Xue Li, Kun Qian +3
In Large Language Model (LLM) inference services, it is challenging to make a parallelism strategy configuration, to efficiently process the requests of variance context lengths. R…