collaborators

5 papers

cs.LG2026

Understanding the Robustness of Distributed Self-Supervised Learning Frameworks Against Non-IID Data

Xuanyu Chen, Nan Yang, Shuai Wang +1

Recent research has introduced distributed self-supervised learning (D-SSL) approaches to leverage vast amounts of unlabeled decentralized data. However, D-SSL faces the critical c…

cs.CL2026

Mechanism-Driven Monitors for Preemptive Detection of LLM Training Instability

Ruixuan Huang, Yipei Wang, Wenyi Fang +7

Frontier large language model training consumes massive accelerator fleets and long wall-clock computation, making stability failures costly when they occur. After a numerical or a…

cs.CR2026

RepetitionCurse: Measuring and Understanding Router Imbalance in Mixture-of-Experts LLMs under DoS Stress

Ruixuan Huang, Qingyue Wang, Hantao Huang +4

Mixture-of-Experts architectures have become the standard for scaling large language models due to their superior parameter efficiency. To accommodate the growing number of experts…

cs.CL2025

SQ-format: A Unified Sparse-Quantized Hardware-friendly Data Format for LLMs

Ruixuan Huang, Hao Zeng, Hantao Huang +4

Post-training quantization (PTQ) plays a crucial role in the democratization of large language models (LLMs). However, existing low-bit quantization and sparsification techniques a…

cs.LG2025

Scaling Law Analysis in Federated Learning: How to Select the Optimal Model Size?

Xuanyu Chen, Nan Yang, Shuai Wang +1

The recent success of large language models (LLMs) has sparked a growing interest in training large-scale models. As the model size continues to scale, concerns are growing about t…