3 papers
cs.LG2026
Parallelism Strategy Chaining for Fast Training Convergence
Minchul Kang, Changyong Shin, Younghun Go +4
Selecting a parallelism strategy - the configuration of data, tensor, and pipeline parallelism degrees together with micro- and global-batch sizes - largely determines the training…
cs.AR2026
Unified KV Pooling to Accelerate Long-Context LLM Serving
Minchul Kang, Changyong Shin, Jinwoo Jeong +6
Long-context LLM serving requires offloading KV caches to host-memory and SSDs, but existing mechanisms are not designed for such long contexts. We observe significant inefficienci…
cs.LG2026
Training Time Prediction for Mixed Precision-based Distributed Training
Minchul Kang, Changyong Shin, Jinwoo Jeong +5
Accurate prediction of training time in distributed deep learning is crucial for resource allocation, cost estimation, and job scheduling. We observe that the floating-point precis…