most citedLoRA meets Riemannion: Muon Optimizer for Parametrization-independent Low-Rank Adapters

1 citations · 1 across the 6 of their papers we have counts for

collaborators

7 papers

cs.LG2026

Controlling Refusal Behavior of LLMs via Stiefel-Constrained Rotation Steering

Kirill Bunin, Dmitry Bylinkin, Vladimir Aletov +3

Activation steering has emerged as a lightweight approach for controlling model refusal at inference time. A growing line of research explores trainable rotations of activations to…

cs.LG2026

Leveraging Association Context Retrieval in Knowledge Edit- ing to Build White-Box Attacks on LLMs

Roman Maksimov, Vladimir Aletov, Vladimir Solodkin +3

As large language models (LLMs) are granted increasing autonomy, it is essential to investigate methods that can induce unsafe behavior. We propose a novel white-box attack inspire…

cs.LG2026

Rethinking the Role of Tensor Decompositions in Post-Training LLM Compression

Artur Zagitov, Alexander Miasnikov, Maxim Krutikov +5

Post-training compression is essential for deploying large language models (LLMs) under tight resource constraints. Tensor decompositions have emerged as a promising direction, off…

cs.LG2026

Tensorion: A Tensor-Aware Generalization of the Muon Optimizer

Vladimir Bogachev, Vladimir Aletov, Alexander Molozhavenko +2

Common first-order optimizers, such as Adam, implicitly treat each parameter block as an unstructured vector, which disregards the multilinear weight structure present in many mode…

cs.LG2026

Scalable Knowledge Editing for Mixture-of-Experts LLMs via Tensor-Structured Updates

Roman Maksimov, Vladimir Aletov, Dmitry Bylinkin +3

Knowledge editing (KE) provides a lightweight alternative to repeated fine-tuning of LLMs. However, most existing KE methods target dense feed-forward layers, while modern LLMs inc…

cs.LG20251 cited

LoRA meets Riemannion: Muon Optimizer for Parametrization-independent Low-Rank Adapters

Vladimir Bogachev, Vladimir Aletov, Alexander Molozhavenko +4

This work presents a novel, fully Riemannian framework for Low-Rank Adaptation (LoRA) that geometrically treats low-rank adapters by optimizing them directly on the fixed-rank mani…