2 papers
cs.NI2026
HCCL: Collective Communication for Meta Training and Inference Accelerators
Wesley Bland, Tiago Antunes, Lars Paul Huse +63
We present HCCL, a collective communication library co-designed with Meta's MTIA 300 accelerator, the first Meta chip to integrate backend networking directly on chip package. MTIA…
cs.DC2026
Training LLMs with Fault Tolerant HSDP on 100,000 GPUs
Omkar Salpekar, Rohan Varma, Kenny Yu +20
Large-scale training systems typically use synchronous training, requiring all GPUs to be healthy simultaneously. In our experience training on O(100K) GPUs, synchronous training r…