2 papers
cs.LG2025
SuperOffload: Unleashing the Power of Large-Scale LLM Training on Superchips
Xinyu Lian, Masahiro Tanaka, Olatunji Ruwase +1
The emergence of Superchips represents a significant advancement in next-generation AI hardware. These Superchips employ a tightly coupled heterogeneous architecture that integrate…
cs.DC2025
Universal Checkpointing: A Flexible and Efficient Distributed Checkpointing System for Large-Scale DNN Training with Reconfigurable Parallelis
Xinyu Lian, Sam Ade Jacobs, Lev Kurilenko +4
Deep neural network (DNN) training continues to scale rapidly in terms of model size, data volume, and sequence length, to the point where multiple machines are required to fit lar…