1 paper
Zichuan Wang, Huizheng Wang, Yuheng Xiao +6
Large language models (LLMs) are increasingly used in prefill-only workloads, where end-to-end latency is dominated by the prefill phase. For long-context prefill, communication ov…