5 citations · 7 across the 4 of their papers we have counts for
5 papers
Linear Interpolation In Parameter Space is Good Enough for Fine-Tuned Language Models
Mark Rofin, Nikita Balagansky, Daniil Gavrilov
The simplest way to obtain continuous interpolation between two points in high dimensional space is to draw a line between them. While previous works focused on the general connect…
FastRPB: a Scalable Relative Positional Encoding for Long Sequence Tasks
Maksim Zubkov, Daniil Gavrilov
Transformers achieve remarkable performance in various domains, including NLP, CV, audio processing, and graph analysis. However, they do not scale well on long sequence tasks due…
Implicit Unlikelihood Training: Improving Neural Text Generation with Reinforcement Learning
Evgeny Lagutin, Daniil Gavrilov, Pavel Kalaidin
Likelihood training and maximization-based decoding result in dull and repetitive generated texts even when using powerful language models (Holtzman et al., 2019). Adding a loss fu…
Weight Squeezing: Reparameterization for Knowledge Transfer and Model Compression
Artem Chumachenko, Daniil Gavrilov, Nikita Balagansky +1
In this work, we present a novel approach for simultaneous knowledge transfer and model compression called Weight Squeezing. With this method, we perform knowledge transfer from a…
Self-Attentive Model for Headline Generation
Daniil Gavrilov, Pavel Kalaidin, Valentin Malykh
Headline generation is a special type of text summarization task. While the amount of available training data for this task is almost unlimited, it still remains challenging, as le…