From the 1 of 7 linked papers with an AI index.
5 papers · 1 filter
Muse: Representation Geometry of Muon Beyond Normalized Momentum
Da Chang, Qiankun Shi, Lvgang Zhang +4
The paper investigates how the choice of matrix representation influences Muon-style optimizers, proposes the Muse family of optimizers that keep the same momentum and Newton–Schul…
A Note on Stability for Orthogonalized Matrix Momentum with Client Sampling
Da Chang, Qiankun Shi, Lvgang Zhang +2
We study finite-sample generalization for a client-sampled distributed optimization scheme with matrix-valued parameters and orthogonalized momentum updates. The central quantity i…
MuonEq: Balancing Before Orthogonalization with Lightweight Equilibration
Da Chang, Qiankun Shi, Lvgang Zhang +5
Orthogonalized-update optimizers such as Muon improve training of matrix-valued parameters, but existing extensions typically either rescale updates after orthogonalization or use…
When Does Value-Aware KV Eviction Help? A Fixed-Contract Diagnostic for Non-Monotone Cache Compression
Ruijie Zhang, Haozhe Liang, Da Chang +4
Long-context LLM inference is bottlenecked by the memory and bandwidth cost of reading large KV caches during decoding. KV compression reduces this cost by keeping only part of the…
AlphaAdam:Asynchronous Masked Optimization with Dynamic Alpha for Selective Updates
Da Chang, Yu Li, Ganzhao Yuan
In the training of large language models (LLMs), updating parameters more efficiently and stably has always been an important challenge. To achieve efficient parameter updates, exi…