1 paper
Ruokai Yin, Sattwik Deb Mishra, Xuan Zuo +3
Distributed LLM inference requires careful coordination of parallelization strategies across hundreds to thousands of NPUs to meet production SLOs. Current systems like Megatron-LM…