collaborators

5 papers

cs.LG2026

Dissecting Outlier Dynamics in LLM NVFP4 Pretraining

Peijie Dong, Ruibo Fan, Yuechen Tao +11

Training large language models using 4-bit arithmetic enhances throughput and memory efficiency. Yet, the limited dynamic range of FP4 increases sensitivity to outliers. While NVFP…

cs.DC2025

Diving into 3D Parallelism with Heterogeneous Spot Instance GPUs: Design and Implications

Yuxiao Wang, Yuedong Xu, Qingyang Duan +4

The rapid growth of large language models (LLMs) and the continuous release of new GPU products have significantly increased the demand for distributed training across heterogeneou…

cs.LG2025

EDGC: Entropy-driven Dynamic Gradient Compression for Efficient LLM Training

Qingao Yi, Jiaang Duan, Hanwen Hu +10

Training large language models (LLMs) poses significant challenges regarding computational resources and memory capacity. Although distributed training techniques help mitigate the…

physics.optics2025

Optical Computation-in-Communication enables low-latency, high-fidelity perception in telesurgery

Rui Yang, Jiaming Hu, Jian-Qing Zheng +12

Artificial intelligence (AI) holds significant promise for enhancing intraoperative perception and decision-making in telesurgery, where physical separation impairs sensory feedbac…

cs.DC2025

GFS: A Preemption-aware Scheduling Framework for GPU Clusters with Predictive Spot Instance Management

Jiaang Duan, Shenglin Xu, Shiyou Qian +15

The surge in large language models (LLMs) has fundamentally reshaped the landscape of GPU usage patterns, creating an urgent need for more efficient management strategies. While cl…