collaborators

5 papers

cs.LG2026

MALT: Lightweight Curvature-Aware Muon via Diagonal Preconditioning

Tongle Wu, Huanyu Dong, Ying Sun +1

Muon has recently emerged as a promising alternative to AdamW for language model pretraining by orthogonalizing momentum matrices using Newton-Schulz iterations. Although Muon miti…

cs.LG2026

Zeta: Dual Whitening for Matrix Optimization via Coordinate-Adaptive Preconditioning

Kaiwen Chen, Shuhai Zhang, Zimo Liu +7

Large-scale neural network training increasingly relies on matrix-aware optimizers that exploit the structure of weight parameters beyond element-wise adaptation. However, existing…

cs.LG2026

FedKRSO: Communication and Memory Efficient Federated Fine-Tuning of Large Language Models

Guohao Yang, Tongle Wu, Yuanxiong Guo +2

Fine-tuning is essential to adapt general-purpose large language models (LLMs) to domain-specific tasks. As a privacy-preserving framework to leverage decentralized data for collab…

cs.LG2024

The Effectiveness of Local Updates for Decentralized Learning under Data Heterogeneity

Tongle Wu, Zhize Li, Ying Sun

We revisit two fundamental decentralized optimization methods, Decentralized Gradient Tracking (DGT) and Decentralized Gradient Descent (DGD), with multiple local updates. We consi…

cs.LG2024

Non-Convex Tensor Recovery from Local Measurements

Tongle Wu, Ying Sun, Jicong Fan

Motivated by the settings where sensing the entire tensor is infeasible, this paper proposes a novel tensor compressed sensing model, where measurements are only obtained from sens…