2 papers
cs.LG2025
Liger Kernel: Efficient Triton Kernels for LLM Training
Pin-Lun Hsu, Yun Dai, Vignesh Kothapalli +7
Training Large Language Models (LLMs) efficiently at scale presents a formidable challenge, driven by their ever-increasing computational demands and the need for enhanced performa…
cs.LG2024
Enhancing Stability for Large Language Models Training in Constrained Bandwidth Networks
Yun Dai, Tejas Dharamsi, Byron Hsu +2
Training extremely large language models (LLMs) with billions of parameters is a computationally intensive task that pushes the limits of current data parallel training systems. Wh…