5 papers
Tensorion: A Tensor-Aware Generalization of the Muon Optimizer
Vladimir Bogachev, Vladimir Aletov, Alexander Molozhavenko +2
Common first-order optimizers, such as Adam, implicitly treat each parameter block as an unstructured vector, which disregards the multilinear weight structure present in many mode…
FiMMIA: scaling semantic perturbation-based membership inference across modalities
Anton Emelyanov, Sergei Kudriashov, Alena Fenogenova
Membership Inference Attacks (MIAs) aim to determine whether a specific data point was included in the training set of a target model. Although there are have been numerous methods…
Pay Attention to Attention Distribution: A New Local Lipschitz Bound for Transformers
Nikolay Yudin, Alexander Gaponov, Sergei Kudriashov +1
We introduce a novel upper bound on the local Lipschitz constant of the dot-product self-attention block showing its dependence on the attention map distributions. The proposed bou…
The more polypersonal the better -- a short look on space geometry of fine-tuned layers
Sergei Kudriashov, Veronika Zykova, Angelina Stepanova +2
The interpretation of deep learning models is a rapidly growing field, with particular interest in language models. There are various approaches to this task, including training si…
Shrink the longest: improving latent space isotropy with symplicial geometry
Sergei Kudriashov, Olesya Karpik, Eduard Klyshinsky
Although transformer-based models have been dominating the field of deep learning, various studies of their embedding space have shown that they suffer from "representation degener…