1 citations · 1 across the 4 of their papers we have counts for
Showing cs.DCShow all
2 papers · 1 filter
cs.DC2026
Bandwidth-Aware and Cost-Efficient Pipeline Parallel Scheduling in Geo-Distributed LLM Training
Han Zhang, Jianchun Liu, Hongli Xu
The rapid evolution of large language models (LLMs) has made geographically distributed training necessary due to GPU scarcity within a single cloud region. In such cross-region se…
cs.DC2025
OptPipe: Memory- and Scheduling-Optimized Pipeline Parallelism for LLM Training
Hongpei Li, Han Zhang, Huikang Liu +2
Pipeline parallelism (PP) has become a standard technique for scaling large language model (LLM) training across multiple devices. However, despite recent progress in reducing memo…