2 citations · 3 across the 2 of their papers we have counts for
Showing cs.DCShow all
2 papers · 1 filter
cs.DC2025★ 1 cited
Understanding LLM Checkpoint/Restore I/O Strategies and Patterns
Mikaila J. Gossman, Avinash Maurya, Bogdan Nicolae +1
As LLMs and foundation models scale, checkpoint/restore has become a critical pattern for training and inference. With 3D parallelism (tensor, pipeline, data), checkpointing involv…
cs.DC2025★ 2 cited
MLP-Offload: Multi-Level, Multi-Path Offloading for LLM Pre-training to Break the GPU Memory Wall
Avinash Maurya, M. Mustafa Rafique, Franck Cappello +1
Training LLMs larger than the aggregated memory of multiple GPUs is increasingly necessary due to the faster growth of LLM sizes compared to GPU memory. To this end, multi-tier hos…