10 papers
Support Before Frequency in Discrete Diffusion
Adrian Müller, Antoine Gonon, Zebang Shen +2
Discrete diffusion models are increasingly competitive for language modeling, yet it remains unclear how their denoising objectives organize learning. Although these objectives tar…
Revisiting DAgger in the Era of LLM-Agents
Changhao Li, Rushi Qiang, Jiawei Huang +4
Long-horizon LM agents learn from multi-turn interaction, where a single early mistake can alter the subsequent state distribution and derail the whole trajectory. Existing recipes…
Muown: Row-Norm Control for Muon Optimization
Kai Lion, Florian Hübler, Bingcong Li +2
Muon has emerged as a strong competitor to AdamW for language model pre-training, yet its behavior at scale is sensitive to weight decay. Recent work has observed that, for Muon wi…
Select-then-differentiate: Solving Bilevel Optimization with Manifold Lower-level Solution Sets
Saeed Masiha, Zebang Shen, Negar Kiyavash +1
We study optimistic bilevel optimization when the lower-level problem has a non-isolated manifold of minimizers. In this setting, the hyper-objective may be non-differentiable beca…
On the Connectedness of Sublevel Sets in Invex Optimization
Vinzenz Thoma, Zebang Shen, Niao He
Understanding the topology of sublevel sets yields crucial insights into the optimization landscape of non-convex functions. If sublevel sets are connected, local search algorithms…
Manifold Generalization Provably Proceeds Memorization in Diffusion Models
Zebang Shen, Ya-Ping Hsieh, Niao He
Diffusion models often generate novel samples even when the learned score is only \emph{coarse} -- a phenomenon not accounted for by the standard view of diffusion training as dens…