7 papers
FOUNDv2: Learning Unified User Quantized Tokenizers for User Representation
Chuan He, Yang Chen, Bin Dou +10
User representation learning serves as a fundamental pillar for personalized services on large-scale web platforms. Despite its importance, conventional continuous embedding method…
Schattor: Schatten-family methods for deep learning optimization
Bohao Ma, Junyu Zhang, Chuan He
Modern deep learning optimization features heterogeneous parameter structures, noisy gradients, and highly nonconvex landscapes, posing significant challenges for both algorithm de…
DeMuon: A Decentralized Muon for Matrix Optimization over Graphs
Chuan He, Shuyi Ren, Jingwei Mao +1
In this paper, we propose DeMuon, a method for decentralized matrix optimization over a given communication topology. DeMuon incorporates matrix orthogonalization via Newton-Schulz…
Low-rank Orthogonalization for Large-scale Matrix Optimization with Applications to Foundation Model Training
Chuan He, Zhanwang Deng, Zhaosong Lu
Neural network (NN) training is inherently a large-scale matrix optimization problem, yet the matrix structure of NN parameters has long been overlooked. Recently, the optimizer Mu…
Stochastic interior-point methods for smooth conic optimization with applications
Chuan He, Zhanwang Deng
Conic optimization plays a crucial role in many machine learning (ML) problems. However, practical algorithms for conic constrained ML problems with large datasets are often limite…
Faster stochastic cubic regularized Newton methods with momentum
Yiming Yang, Chuan He, Xiao Wang +1
Cubic regularized Newton (CRN) methods have attracted signiffcant research interest because they offer stronger solution guarantees and lower iteration complexity. With the rise of…