4 papers
MALT: Lightweight Curvature-Aware Muon via Diagonal Preconditioning
Tongle Wu, Huanyu Dong, Ying Sun +1
Muon has recently emerged as a promising alternative to AdamW for language model pretraining by orthogonalizing momentum matrices using Newton-Schulz iterations. Although Muon miti…
FedKRSO: Communication and Memory Efficient Federated Fine-Tuning of Large Language Models
Guohao Yang, Tongle Wu, Yuanxiong Guo +2
Fine-tuning is essential to adapt general-purpose large language models (LLMs) to domain-specific tasks. As a privacy-preserving framework to leverage decentralized data for collab…
The Effectiveness of Local Updates for Decentralized Learning under Data Heterogeneity
Tongle Wu, Zhize Li, Ying Sun
We revisit two fundamental decentralized optimization methods, Decentralized Gradient Tracking (DGT) and Decentralized Gradient Descent (DGD), with multiple local updates. We consi…
Non-Convex Tensor Recovery from Local Measurements
Tongle Wu, Ying Sun, Jicong Fan
Motivated by the settings where sensing the entire tensor is infeasible, this paper proposes a novel tensor compressed sensing model, where measurements are only obtained from sens…