5 papers
Adam Converges Without Any Modification On Update Rules
Yushun Zhang, Bingran Li, Congliang Chen +2
Adam is the default algorithm for training neural networks, including large language models (LLMs). However, \citet{reddi2019convergence} provided an example that Adam diverges, ra…
Exact Causal Attention with 10% Fewer Operations
Dmitry Rybin, Yushun Zhang, Ding Tian +2
We present Exact Causal Attention (ECA), a Strassen-style algorithm that computes exact Causal Attention using 10\% fewer operations. ECA improves a special class of matrix multipl…
Can Be Faster
Dmitry Rybin, Yushun Zhang, Zhi-Quan Luo
We present RXTX, a new algorithm for computing the product of matrix by its transpose for . RXTX uses fewer multiplications and fe…
Towards Quantifying the Hessian Structure of Neural Networks
Zhaorui Dong, Yushun Zhang, Jianfeng Yao +1
Empirical studies reported that the Hessian matrix of neural networks (NNs) exhibits a near-block-diagonal structure, yet its theoretical foundation remains unclear. In this work,…
Finite Horizon Optimization: Framework and Applications
Yushun Zhang, Dmitry Rybin, Zhi-Quan Luo
In modern engineering scenarios, there is often a strict upper bound on the number of algorithm iterations that can be performed within a given time limit. This raises the question…