Showing cs.DCShow all
2 papers · 1 filter
cs.DC2025
MeCeFO: Enhancing LLM Training Robustness via Fault-Tolerant Optimization
Rizhen Hu, Yutong He, Ran Yan +3
As distributed optimization scales to meet the demands of Large Language Model (LLM) training, hardware failures become increasingly non-negligible. Existing fault-tolerant trainin…
cs.DC2025
HexGen-2: Disaggregated Generative Inference of LLMs in Heterogeneous Environment
Youhe Jiang, Ran Yan, Binhang Yuan
Disaggregating the prefill and decoding phases represents an effective new paradigm for generative inference of large language models (LLM), which eliminates prefill-decoding inter…