18 citations · 23 across the 12 of their papers we have counts for
1 paper · 1 filter
Amy Yang, Jingyi Yang, Aya Ibrahim +6
We present context parallelism for long-context large language model inference, which achieves near-linear scaling for long-context prefill latency with up to 128 H100 GPUs across…