15 papers
LoRA-TSD: Tangent-Space Spectral Descent for LoRA via Muon-Style Updates
Dmitrii Andriianov, Andrey Veprikov, Aleksandr Beznosikov
Low-rank adaptation (LoRA) is the standard way to fine-tune large models, yet when its two factors are trained independently, the update ignores the geometry of the low-rank weight…
Controlling Refusal Behavior of LLMs via Stiefel-Constrained Rotation Steering
Kirill Bunin, Dmitry Bylinkin, Vladimir Aletov +3
Activation steering has emerged as a lightweight approach for controlling model refusal at inference time. A growing line of research explores trainable rotations of activations to…
Leveraging Association Context Retrieval in Knowledge Edit- ing to Build White-Box Attacks on LLMs
Roman Maksimov, Vladimir Aletov, Vladimir Solodkin +3
As large language models (LLMs) are granted increasing autonomy, it is essential to investigate methods that can induce unsafe behavior. We propose a novel white-box attack inspire…
Self-Study Reconsidered: The Hidden Fragility of Learning from Self-Generated QA
Ekaterina Alimaskina, Denis Shveykin, Gleb Molodtsov +3
Language models are increasingly taught from synthetic question--answer (QA) supervision: a model generates questions about a document, answers them from the same text, and the res…
Zero-order Parameter-free Optimization for LMO-based Methods: Novel Approach for Efficient Fine-tuning
Dmitriy Bystrov, Daniil Medyakov, Dmitry Bylinkin +1
Fine-tuning large language models (LLMs) has become a central application of modern optimization, enabling pretrained models to adapt to diverse downstream tasks and domain-specifi…
Analyzing Stream Collapse in Hyper-Connections: From Diagnosis to Mitigation
Ekaterina Alimaskina, Gleb Molodtsov, Aleksandr Beznosikov
Hyper-Connections (HC) replace the single Transformer residual stream with multiple streams, introducing a permutation symmetry over stream indices. We study how this symmetry is r…