2 papers
cs.DC2026
LayerCheck: Adaptive Layer-wise Checkpointing for Large Language Model Post-training
Minqiu Sun, Xin Huang, Luanzheng Guo +3
With the rising computational and monetary costs of training large language models (LLMs), checkpointing---periodically storing model states for recovery---becomes essential for fa…
cs.DC2026
ZOCheck: CPU-Shadow Checkpointing for Zeroth-Order LLM Fine-Tuning
Minqiu Sun, Xin Huang, Luanzheng Guo +3
Zeroth-order (ZO) optimization is an attractive option for memory-efficient LLM fine-tuning, but its fault tolerance remains underexplored. Unlike first-order training, ZO progress…