3 papers
cs.LG2026
MALT: Lightweight Curvature-Aware Muon via Diagonal Preconditioning
Tongle Wu, Huanyu Dong, Ying Sun +1
Muon has recently emerged as a promising alternative to AdamW for language model pretraining by orthogonalizing momentum matrices using Newton-Schulz iterations. Although Muon miti…
cs.LG2026
FedKRSO: Communication and Memory Efficient Federated Fine-Tuning of Large Language Models
Guohao Yang, Tongle Wu, Yuanxiong Guo +2
Fine-tuning is essential to adapt general-purpose large language models (LLMs) to domain-specific tasks. As a privacy-preserving framework to leverage decentralized data for collab…
cs.LG2024
Non-Convex Tensor Recovery from Local Measurements
Tongle Wu, Ying Sun, Jicong Fan
Motivated by the settings where sensing the entire tensor is infeasible, this paper proposes a novel tensor compressed sensing model, where measurements are only obtained from sens…