activity
20242026
collaborators

12 papers

math.OC2026

Chebyshev-Exact Acceleration under Hessian Variation, I: Sine-Jacobi Method

Dmitry Pasechnyuk-Vilensky, Martin Takáč

We study finite-horizon one-gradient realizations with the Chebyshev minimax terminal residual on . Under time-dependent Hessian perturbations, the terminal first variation…

cs.LG2026

Privacy from Symmetry: Orthogonally Equivariant Transformers for LLM Inference

Alexander Yukhimchuk, Andrey Shulga, Mladen Kolar +1

Running large language models locally is often impractical, pushing inference on sensitive text to third-party providers. Split inference partially mitigates this by keeping tokens…

math.OC2026

Decentralized Inexact Cubic Newton Method with Consensus Procedure

Artem Agafonov, Anton Novitskii, Alexander Rogozin +5

Distributed optimization is widely used in large-scale and privacy-preserving machine learning, where each agent stores a local objective and communicates only with its neighbors i…

cs.LG2026

Where Does Warm-Up Come From? Adaptive Scheduling for Norm-Constrained Optimizers

Artem Riabinin, Andrey Veprikov, Arman Bolatov +2

We study adaptive learning rate scheduling for norm-constrained optimizers (e.g., Muon and Lion). We introduce a generalized smoothness assumption under which local curvature decre…

cs.LG2026

Gradient Clipping Beyond Vector Norms: A Spectral Approach for Matrix-Valued Parameters

Alexander Yukhimchuk, Mladen Kolar, Martin Takáč +1

Gradient clipping is a standard safeguard for training neural networks under noisy, heavy-tailed stochastic gradients; yet, most clipping rules treat all parameters as vectors and…

math.OC2026

Muon with Nesterov Momentum: Heavy-Tailed Noise and (Randomized) Inexact Polar Decomposition

Sayantan Choudhury, Xiaoran Cheng, Martin Takáč +2

Most first-order optimizers treat matrix-valued parameters as vectors, ignoring the intrinsic geometry of hidden-layer weights in neural networks. Muon addresses this mismatch by u…