1 citations · 1 across the 2 of their papers we have counts for
2 papers
cs.DC2026
VPP: Virtual Pipeline Parallelism for Efficient Chunked Prefill in Long-Context LLM Inference
Yan Shi, Xiaochao Wang, Jingchun Gao +7
Chunked prefill pipeline parallelism (CPP) is a key technique for LLM inference. However, equal-size chunks exhibit imbalanced latency, as later chunks attend longer prefix KV cach…
cs.LG2025★ 1 cited
AsyncFlow: An Asynchronous Streaming RL Framework for Efficient LLM Post-Training
Zhenyu Han, Ansheng You, Haibo Wang +16
Reinforcement learning (RL) has become a pivotal technology in the post-training phase of large language models (LLMs). Traditional task-collocated RL frameworks suffer from signif…