Showing cs.DCShow all
2 papers · 1 filter
cs.DC2026
SPARe: Stacked Parallelism with Adaptive Reordering for Fault-Tolerant LLM Pretraining Systems with 100k+ GPUs
Jin Lee, Zhonghao Chen, Xuhang He +6
In large-scale LLM pre-training systems with 100k+ GPUs, failures become the norm rather than the exception, and restart costs can dominate wall-clock training time. However, exist…
cs.DC2024
DataStates-LLM: Lazy Asynchronous Checkpointing for Large Language Models
Avinash Maurya, Robert Underwood, M. Mustafa Rafique +2
LLMs have seen rapid adoption in all domains. They need to be trained on high-end high-performance computing (HPC) infrastructures and ingest massive amounts of input data. Unsurpr…