1 paper
Daniel Waddington, Cornel Constantinescu
During the training of Large Language Models (LLMs), tensor data is periodically "checkpointed" to persistent storage to allow recovery of work done in the event of failure. The vo…