1 paper
Bohan Zhao, Yuanhong Wang, Chenglin Liu +6
Recent developments in large language models (LLMs) have introduced new requirements for efficient and robust training. As LLM clusters scale, node failures, lengthy recoveries, an…