5 papers
Tensorion: A Tensor-Aware Generalization of the Muon Optimizer
Vladimir Bogachev, Vladimir Aletov, Alexander Molozhavenko +2
Common first-order optimizers, such as Adam, implicitly treat each parameter block as an unstructured vector, which disregards the multilinear weight structure present in many mode…
Rethinking the Role of Tensor Decompositions in Post-Training LLM Compression
Artur Zagitov, Alexander Miasnikov, Maxim Krutikov +5
Post-training compression is essential for deploying large language models (LLMs) under tight resource constraints. Tensor decompositions have emerged as a promising direction, off…
Scalable Knowledge Editing for Mixture-of-Experts LLMs via Tensor-Structured Updates
Roman Maksimov, Vladimir Aletov, Dmitry Bylinkin +3
Knowledge editing (KE) provides a lightweight alternative to repeated fine-tuning of LLMs. However, most existing KE methods target dense feed-forward layers, while modern LLMs inc…
Bant: Byzantine Antidote via Trial Function and Trust Scores
Gleb Molodtsov, Daniil Medyakov, Sergey Skorik +6
Recent advancements in machine learning have improved performance while also increasing computational demands. While federated and distributed setups address these issues, their st…
LoRA meets Riemannion: Muon Optimizer for Parametrization-independent Low-Rank Adapters
Vladimir Bogachev, Vladimir Aletov, Alexander Molozhavenko +4
This work presents a novel, fully Riemannian framework for Low-Rank Adaptation (LoRA) that geometrically treats low-rank adapters by optimizing them directly on the fixed-rank mani…