2 papers
cs.DC2026
Amoeba: Runtime Tensor Parallel Transformation for LLM Inference Services
Haoyu Chen, Xue Li, Kun Qian +3
In Large Language Model (LLM) inference services, it is challenging to make a parallelism strategy configuration, to efficiently process the requests of variance context lengths. R…
cs.DC2026
EROICA: Online Performance Troubleshooting for Large-scale Model Training
Yu Guan, Zhiyu Yin, Haoyu Chen +11
Troubleshooting performance problems of large model training (LMT) is immensely challenging, due to unprecedented scales of modern GPU clusters, the complexity of software-hardware…