3 papers
cs.LG2026
LoKiFormer: Locality-aware Attention with Decoupled Knowledge Memory for Efficient Large Language Model Pretraining
Qiuwu Chen, Zimo Liu, Yuchen Li +8
Large language models (LLMs) have achieved remarkable breakthroughs across various applications. However, their architectures remain inefficient in pretraining due to two main limi…
cs.LG2026
Zeta: Dual Whitening for Matrix Optimization via Coordinate-Adaptive Preconditioning
Kaiwen Chen, Shuhai Zhang, Zimo Liu +7
Large-scale neural network training increasingly relies on matrix-aware optimizers that exploit the structure of weight parameters beyond element-wise adaptation. However, existing…
math.OC2024
Convergence and Complexity Guarantee for Inexact First-order Riemannian Optimization Algorithms
Yuchen Li, Laura Balzano, Deanna Needell +1
We analyze inexact Riemannian gradient descent (RGD) where Riemannian gradients and retractions are inexactly (and cheaply) computed. Our focus is on understanding when inexact RGD…