2 papers
cs.DC2025
Checkmate: Zero-Overhead Model Checkpointing via Network Gradient Replication
Ankit Bhardwaj, Weiyang Wang, Jeremy Carin +2
This paper presents Checkmate, a system that enables per-iteration checkpointing in DNN training without any training slowdown. The traditional approach to checkpointing requires a…
cs.NI2024
Rail-only: A Low-Cost High-Performance Network for Training LLMs with Trillion Parameters
Weiyang Wang, Manya Ghobadi, Kayvon Shakeri +2
This paper presents a low-cost network architecture for training large language models (LLMs) at hyperscale. We study the optimal parallelization strategy of LLMs and propose a nov…