2 papers
cs.DC2026
HCMS: Head-Chunked Multi-Stream Pipeline for Communication-Computation Overlap in Long-Sequence Parallel Attention
Chao Yuan, Pan Li, Yingnan Sun +1
All-to-all based sequence parallelism methods execute communication and computation strictly in serial when processing medium-long sequences, resulting in hardware resource underut…
cs.CV2026
Accelerating Video Generation Inference with Sequential-Parallel 3D Positional Encoding Using a Global Time Index
Chao Yuan, Pan Li
Diffusion Transformer (DiT)-based video generation models inherently suffer from bottlenecks in long video synthesis and real-time inference, which can be attributed to the use of…