18 papers
LionMuon: Alternating Spectral and Sign Descent for Efficient Training
Arman Bolatov, Artem Riabinin, Nikita Kornilov +6
In large-scale optimization, the cheapness and effectiveness of update steps are the most crucial factors for a successful optimizer. Sign-based optimizers like Lion or Signum prod…
CacheMuon: Using Temporal Preconditioning To Approximate Polar Factor
Bishnu Dev, Sushil Bohara, Martin TakÃ¡Ä +1
Muon is an optimizer that computes updates using the polar factor of the momentum matrix and has shown strong empirical performance across a range of training settings. A key compo…
Value-Gradient Hypothesis of RL for LLMs
Arip Asadulaev, Daniil Ognev, Karim Salta +1
Reinforcement learning substantially improves pretrained language models, but it remains understudied why critic-free methods such as PPO and GRPO work as well as they do, and when…
Tractable Probabilistic Models for Investment Planning
Nicolas M. Cuadrado A., Mohannad Takrouri, JiÅÃ NÄmeÄek +2
Investment planning in power utilities, such as generation and transmission expansion, requires decisions under substantial uncertainty over decade--long horizons for policies, dem…
Byzantine-Robust Optimization under -Smoothness
Arman Bolatov, Samuel Horváth, Martin TakÃ¡Ä +1
We consider distributed optimization under Byzantine attacks in the presence of -smoothness, a generalization of standard -smoothness that captures functions with sta…
LoFT: Low-Rank Adaptation That Behaves Like Full Fine-Tuning
Nurbek Tastan, Stefanos Laskaridis, Martin Takac +2
Large pre-trained models are commonly adapted to downstream tasks using parameter-efficient fine-tuning methods such as Low-Rank Adaptation (LoRA), which injects small trainable lo…