2 citations · 2 across the 4 of their papers we have counts for
Showing math.OCShow all
3 papers · 1 filter
math.OC2026
Symmetry-Compatible Principle for Optimizer Design: Embeddings, LM Heads, SwiGLU MLPs, and MoE Routers
Tim Tsz-Kit Lau, Weijie Su
A striking geometric disparity has long persisted in the practice of deep learning. While modern neural network architectures naturally exhibit rich symmetry and equivariance prope…
math.OC2025
PolarGrad: A Class of Matrix-Gradient Optimizers from a Unifying Preconditioning Perspective
Tim Tsz-Kit Lau, Qi Long, Weijie Su
The ever-growing scale of deep learning models and training data underscores the critical importance of efficient optimization methods. While preconditioned gradient methods such a…
math.OC2018
Global Convergence of Block Coordinate Descent in Deep Learning
Jinshan Zeng, Tim Tsz-Kit Lau, Shaobo Lin +1
Deep learning has aroused extensive attention due to its great empirical success. The efficiency of the block coordinate descent (BCD) methods has been recently demonstrated in dee…