12 papers
Chebyshev-Exact Acceleration under Hessian Variation, I: Sine-Jacobi Method
Dmitry Pasechnyuk-Vilensky, Martin TakáÄ
We study finite-horizon one-gradient realizations with the Chebyshev minimax terminal residual on . Under time-dependent Hessian perturbations, the terminal first variation…
Privacy from Symmetry: Orthogonally Equivariant Transformers for LLM Inference
Alexander Yukhimchuk, Andrey Shulga, Mladen Kolar +1
Running large language models locally is often impractical, pushing inference on sensitive text to third-party providers. Split inference partially mitigates this by keeping tokens…
Decentralized Inexact Cubic Newton Method with Consensus Procedure
Artem Agafonov, Anton Novitskii, Alexander Rogozin +5
Distributed optimization is widely used in large-scale and privacy-preserving machine learning, where each agent stores a local objective and communicates only with its neighbors i…
Where Does Warm-Up Come From? Adaptive Scheduling for Norm-Constrained Optimizers
Artem Riabinin, Andrey Veprikov, Arman Bolatov +2
We study adaptive learning rate scheduling for norm-constrained optimizers (e.g., Muon and Lion). We introduce a generalized smoothness assumption under which local curvature decre…
Gradient Clipping Beyond Vector Norms: A Spectral Approach for Matrix-Valued Parameters
Alexander Yukhimchuk, Mladen Kolar, Martin TakÃ¡Ä +1
Gradient clipping is a standard safeguard for training neural networks under noisy, heavy-tailed stochastic gradients; yet, most clipping rules treat all parameters as vectors and…
Muon with Nesterov Momentum: Heavy-Tailed Noise and (Randomized) Inexact Polar Decomposition
Sayantan Choudhury, Xiaoran Cheng, Martin TakÃ¡Ä +2
Most first-order optimizers treat matrix-valued parameters as vectors, ignoring the intrinsic geometry of hidden-layer weights in neural networks. Muon addresses this mismatch by u…