24 citations · 30 across the 3 of their papers we have counts for
3 papers
cs.DC2025
TileLink: Generating Efficient Compute-Communication Overlapping Kernels using Tile-Centric Primitives
Size Zheng, Jin Fang, Xuegui Zheng +9
Large deep learning models have achieved state-of-the-art performance in a wide range of tasks. These models often necessitate distributed systems for efficient training and infere…
cs.LG2024★ 24 cited
MegaScale: Scaling Large Language Model Training to More Than 10,000 GPUs
Ziheng Jiang, Haibin Lin, Yinmin Zhong +29
We present the design, implementation and engineering experience in building and deploying MegaScale, a production system for training large language models (LLMs) at the scale of…
cs.NI2023★ 6 cited
Collie: Finding Performance Anomalies in RDMA Subsystems
Xinhao Kong, Yibo Zhu, Huaping Zhou +4
High-speed RDMA networks are getting rapidly adopted in the industry for their low latency and reduced CPU overheads. To verify that RDMA can be used in production, system administ…