2 citations
- Nvidia (United Kingdom)GB6 papers
- Harvard UniversityUS2 papers
- Stanford UniversityUS2 papers
- Abterra Biosciences (United States)US1 paper
- Air Water (Japan)JP1 paper
- Argonne National LaboratoryUS1 paper
- Brigham and Women's HospitalUS1 paper
- California Institute of TechnologyUS1 paper
- Children's NationalUS1 paper
- ETH ZurichCH1 paper
- Fermi National Accelerator LaboratoryUS1 paper
- George Washington UniversityUS1 paper
Showing cs.DCShow all
2 papers · 1 filter
cs.DC2026
NCCL EP: Towards a Unified Expert Parallel Communication API for NCCL
Amos Goldman, Nimrod Boker, Maayan Sheraizin +15
Mixture-of-Experts (MoE) architectures have become essential for scaling large language models, driving the development of specialized device-initiated communication libraries such…
cs.DC2026
Parallelizing the Approximate Minimum Degree Ordering Algorithm: Strategies and Evaluation
Yen-Hsiang Chang, Aydın Buluç, James Demmel
The approximate minimum degree algorithm is widely used before numerical factorization to reduce fill-in for sparse matrices. While considerable attention has been given to the num…