collaborators

6 papers

cs.DC2025

Oases: Efficient Large-Scale Model Training on Commodity Servers via Overlapped and Automated Tensor Model Parallelism

Shengwei Li, Zhiquan Lai, Dongsheng Li +5

Deep learning is experiencing a rise in large-scale models. Training large-scale models is costly, prompting researchers to train large-scale models on commodity servers that more…

cs.LG2025

Towards Understanding the Generalizability of Delayed Stochastic Gradient Descent

Xiaoge Deng, Li Shen, Shengwei Li +3

Stochastic gradient descent (SGD) performed in an asynchronous manner plays a crucial role in training large-scale machine learning models. However, the generalization performance…

cs.LG2024

Sharpness-Aware Minimization with Adaptive Regularization for Training Deep Neural Networks

Jinping Zou, Xiaoge Deng, Tao Sun

Sharpness-Aware Minimization (SAM) has proven highly effective in improving model generalization in machine learning tasks. However, SAM employs a fixed hyperparameter associated w…

cs.LG2024

Federated Prediction-Powered Inference from Decentralized Data

Ping Luo, Xiaoge Deng, Ziqing Wen +2

In various domains, the increasing application of machine learning allows researchers to access inexpensive predictive data, which can be utilized as auxiliary data for statistical…

cs.DC2024

Communication-Efficient Distributed Learning via Sparse and Adaptive Stochastic Gradient

Xiaoge Deng, Dongsheng Li, Tao Sun +1

Gradient-based optimization methods implemented on distributed computing architectures are increasingly used to tackle large-scale machine learning applications. A key bottleneck i…

cs.LG2024

Score-based Generative Models with Adaptive Momentum

Ziqing Wen, Xiaoge Deng, Ping Luo +2

Score-based generative models have demonstrated significant practical success in data-generating tasks. The models establish a diffusion process that perturbs the ground truth data…