1 citations · 1 across the 7 of their papers we have counts for
1 paper · 1 filter
Amy Yang, Jingyi Yang, Aya Ibrahim +6
We present context parallelism for long-context large language model inference, which achieves near-linear scaling for long-context prefill latency with up to 128 H100 GPUs across…