From the 23 of 585 papers with an AI index.
884 citations
- University of Chinese Academy of SciencesCN190 papers
- Peking UniversityCN165 papers
- University of Science and Technology of ChinaCN164 papers
- Istituto Nazionale di Fisica Nucleare, Laboratori Nazionali di FrascatiIT144 papers
- Université Paris-SaclayFR130 papers
- The Ohio State UniversityUS128 papers
- Jagiellonian UniversityPL127 papers
- Rutherford Appleton LaboratoryGB127 papers
- Istituto Nazionale di Fisica Nucleare, Sezione di BolognaIT124 papers
- AGH University of KrakowPL122 papers
- South China Normal UniversityCN121 papers
- Istituto Nazionale di Fisica Nucleare, Sezione di Roma IIT120 papers
8 papers · 1 filter
Collective Communication for Distributed LLM Systems: Planning, Runtime Adaptation, and Computation Coordination
Xuebin Song, Menghao Zhang, Yuezheng Liu +5
Distributed large language model (LLM) systems increasingly rely on collective communication primitives such as AllReduce (AR), ReduceScatter (RS), AllGather (AG), and AlltoAll (A2…
Application-Driven Architecture Exploration for Cross-Layer Heterogeneous Systems
Yuchen Fan, Minghong Sun, Jikui Ma +19
AI and HPC infrastructure increasingly serves workload portfolios that combine dense tensor computation, sparse kernels, large memory footprints, and communication-intensive collec…
Odin: Primitive-Level Synchronization for Distributed Point-Based Neural Rendering
Zhenxiang Ma, Zeyu He, Yuanzhen Zhou +6
Point-based neural rendering (PBNR) represents 3D scenes as explicit, trainable primitives and underpins high-quality reconstruction and emerging embodied AI and world-model pipeli…
Don't Predict, Prioritize: Rethinking GPU Reliability Assessment
Difeng Ma, Changhua Pei, Yuanwei Lu +7
The paper proposes HeaRank, a learning-to-rank framework that ranks GPU nodes by their relative failure risk instead of predicting exact failure times, showing improved detection o…
An Efficient, Reliable and Observable Collective Communication Library in Large-scale GPU Training Clusters
Mingjun Zhang, Xiaohe Hu, Menghao Zhang +21
Large-scale LLM training requires collective communication libraries to exchange data among distributed GPUs. As a company dedicated to building and operating large-scale GPU train…
A Readiness-Driven Runtime for Pipeline-Parallel Training under Runtime Variability
Ruitao Liu, Xinyang Tian, Shuo Chen +4
Pipeline parallelism is a key technique for scaling large-model training, but modern workloads exhibit runtime variability in computation and communication. Existing pipeline syste…