3 papers
cs.DC2026
Orbax: Distributed Checkpointing with JAX
Colin Gaffney, Shutong Li, Daniel Ng +13
In a landscape of high-performance distributed ML systems, JAX has emerged as a framework of choice. However, JAX's modular design philosophy leaves it without a standardized check…
cs.DC2025
Efficient Distributed MLLM Training with Cornstarch
Insu Jang, Runyu Lu, Nikhil Bansal +2
Multimodal large language models (MLLMs) extend the capabilities of large language models (LLMs) by combining heterogeneous model architectures to handle diverse modalities like im…
cs.LG2023
Reducing Energy Bloat in Large Model Training
Jae-Won Chung, Yile Gu, Insu Jang +3
Training large AI models on numerous GPUs consumes a massive amount of energy, making power delivery one of the largest limiting factors in building and operating datacenters for A…