1 paper
Yan Liang, Youhe Jiang, Ran Yan +3
Long-context training of large language models (LLMs) is commonly distributed with Context Parallelism (CP) and Head Parallelism (HP), but existing training systems largely assume…