6 papers
FedACT: Federated Adaptive Coordinate Trust Modulation for Robust Transformer Training under Data Heterogeneity
Shuai Li, Qinglin Wang, Ping Luo +8
Federated Transformer training increasingly relies on local AdamW, whose adaptive updates can provide much stronger local progress than SGD-based training. However, under heterogen…
FOAM: Blocked State Folding for Memory-Efficient LLM Training
Ziqing Wen, Jiahuan Wang, Ping Luo +2
Large language models (LLMs) have demonstrated remarkable performance due to their large parameter counts and extensive training data. However, their scale leads to significant mem…
Revealing Modular Gradient Noise Imbalance in LLMs: Calibrating Adam via Signal-to-Noise Ratio
Ziqing Wen, Zhouyang Liu, Jiahuan Wang +4
The impressive performance of large language models (LLMs) arises from their massive scale and heterogeneous module composition. However, this structural heterogeneity introduces a…
Stability and Generalization for Decentralized Markov SGD
Jiahuan Wang, Ziqing Wen, Ping Luo +2
Stochastic gradient methods are central to large-scale learning, yet their generalization theory typically relies on independent sampling assumptions. In many practical application…
GWT: Scalable Optimizer State Compression for Large Language Model Training
Ziqing Wen, Ping Luo, Jiahuan Wang +3
Large Language Models (LLMs) have demonstrated exceptional capabilities across diverse natural language processing benchmarks. However, the escalating scale of model parameters imp…
Local Gradient Regulation Stabilizes Federated Learning under Client Heterogeneity
Ping Luo, Jiahuan Wang, Ziqing Wen +2
Federated learning (FL) enables collaborative model training across distributed clients without sharing raw data, yet its stability is fundamentally challenged by statistical heter…