9 papers
Muown Implicitly Performs Angular Step-size Decay
Florian Hübler, Kai Lion, Antonio Orvieto +1
Matrix-aware optimizers such as Muon and Muown have recently shown strong empirical performance for pre-training Transformers. In particular, Muown separates each weight matrix int…
When Scores Learn Geometry: Rate Separations under the Manifold Hypothesis
Xiang Li, Zebang Shen, Ya-Ping Hsieh +1
Score-based methods, such as diffusion models and Bayesian inverse problems, are often interpreted as learning the data distribution in the low-noise limit (). In this wor…
Landing with the Score: Riemannian Optimization through Denoising
Andrey Kharitenko, Zebang Shen, Riccardo de Santi +2
Under the data manifold hypothesis, high-dimensional data are concentrated near a low-dimensional manifold. We study the problem of Riemannian optimization over such manifolds when…
ANCRe: Adaptive Neural Connection Reassignment for Efficient Depth Scaling
Yilang Zhang, Bingcong Li, Niao He +1
Scaling network depth has been a central driver behind the success of modern foundation models, yet recent investigations suggest that deep layers are often underutilized. This pap…
Flow Density Control: Generative Optimization Beyond Entropy-Regularized Fine-Tuning
Riccardo De Santi, Marin Vlastelica, Ya-Ping Hsieh +3
Adapting large-scale foundation flow and diffusion generative models to optimize task-specific objectives while preserving prior information is crucial for real-world applications…
Optical Computing with Spectrally Multiplexed Features in Complex Media
Xue Dong, Kai Lion, Fei Xia +4
Artificial intelligence (AI) has rapidly evolved into a critical technology; however, electrical hardware struggles to keep pace with the exponential growth of AI models. Free spac…