collaborators

6 papers

cs.LG2026

Muse: Representation Geometry of Muon Beyond Normalized Momentum

Da Chang, Qiankun Shi, Lvgang Zhang +4

Muon-style optimizers apply a polar map to matrix momentum, but their updates also depend on the representation of each parameter block before orthogonalization. We study this repr…

cs.LG2026

A Note on Stability for Orthogonalized Matrix Momentum with Client Sampling

Da Chang, Qiankun Shi, Lvgang Zhang +2

We study finite-sample generalization for a client-sampled distributed optimization scheme with matrix-valued parameters and orthogonalized momentum updates. The central quantity i…

cs.LG2026

When Does Value-Aware KV Eviction Help? A Fixed-Contract Diagnostic for Non-Monotone Cache Compression

Ruijie Zhang, Haozhe Liang, Da Chang +4

Long-context LLM inference is bottlenecked by the memory and bandwidth cost of reading large KV caches during decoding. KV compression reduces this cost by keeping only part of the…

cs.LG2026

MuonEq: Balancing Before Orthogonalization with Lightweight Equilibration

Da Chang, Qiankun Shi, Lvgang Zhang +5

Orthogonalized-update optimizers such as Muon improve training of matrix-valued parameters, but existing extensions typically either rescale updates after orthogonalization or use…

cs.CV2025

Encoding Structural Constraints into Segment Anything Models via Probabilistic Graphical Models

Yu Li, Da Chang, Xi Xiao

While the Segment Anything Model (SAM) has achieved remarkable success in image segmentation, its direct application to medical imaging remains hindered by fundamental challenges,…

cs.LG2025

AlphaAdam:Asynchronous Masked Optimization with Dynamic Alpha for Selective Updates

Da Chang, Yu Li, Ganzhao Yuan

In the training of large language models (LLMs), updating parameters more efficiently and stably has always been an important challenge. To achieve efficient parameter updates, exi…