2 papers
cs.DC2025
TTrace: Lightweight Error Checking and Diagnosis for Distributed Training
Haitian Jiang, Shaowei Zhu, Zhen Zhang +5
Distributed training is essential for scaling the training of large neural network models, such as large language models (LLMs), across thousands of GPUs. However, the complexity o…
cs.DC2024
PipeFill: Using GPUs During Bubbles in Pipeline-parallel LLM Training
Daiyaan Arfeen, Zhen Zhang, Xinwei Fu +2
Training Deep Neural Networks (DNNs) with billions of parameters generally involves pipeline-parallel (PP) execution. Unfortunately, PP model training can use GPUs inefficiently, e…