works on

From the 23 of 585 papers with an AI index.

most citedDeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

884 citations

Showing cs.DCShow all

8 papers · 1 filter

cs.DC2026

Collective Communication for Distributed LLM Systems: Planning, Runtime Adaptation, and Computation Coordination

Xuebin Song, Menghao Zhang, Yuezheng Liu +5

Distributed large language model (LLM) systems increasingly rely on collective communication primitives such as AllReduce (AR), ReduceScatter (RS), AllGather (AG), and AlltoAll (A2…

cs.DC2026

Application-Driven Architecture Exploration for Cross-Layer Heterogeneous Systems

Yuchen Fan, Minghong Sun, Jikui Ma +19

AI and HPC infrastructure increasingly serves workload portfolios that combine dense tensor computation, sparse kernels, large memory footprints, and communication-intensive collec…

cs.DC2026

Odin: Primitive-Level Synchronization for Distributed Point-Based Neural Rendering

Zhenxiang Ma, Zeyu He, Yuanzhen Zhou +6

Point-based neural rendering (PBNR) represents 3D scenes as explicit, trainable primitives and underpins high-quality reconstruction and emerging embodied AI and world-model pipeli…

cs.DC2026

Don't Predict, Prioritize: Rethinking GPU Reliability Assessment

Difeng Ma, Changhua Pei, Yuanwei Lu +7

The paper proposes HeaRank, a learning-to-rank framework that ranks GPU nodes by their relative failure risk instead of predicting exact failure times, showing improved detection o…

cs.DC2026

An Efficient, Reliable and Observable Collective Communication Library in Large-scale GPU Training Clusters

Mingjun Zhang, Xiaohe Hu, Menghao Zhang +21

Large-scale LLM training requires collective communication libraries to exchange data among distributed GPUs. As a company dedicated to building and operating large-scale GPU train…

cs.DC2026

A Readiness-Driven Runtime for Pipeline-Parallel Training under Runtime Variability

Ruitao Liu, Xinyang Tian, Shuo Chen +4

Pipeline parallelism is a key technique for scaling large-model training, but modern workloads exhibit runtime variability in computation and communication. Existing pipeline syste…