1 paper
Aojie Jiang, Kang Zhu, Zhiheng Zhang +4
Tensor parallelism (TP) has become a key technique for latency-sensitive LLM inference, but it introduces frequent, tightly synchronized All-Reduce operations that lie directly on…