330 citations
- Skolkovo Institute of Science and TechnologyRU7 papers
- National Research University Higher School of EconomicsRU6 papers
- Lomonosov Moscow State UniversityRU5 papers
- Moscow Institute of Physics and TechnologyRU5 papers
- ITMO UniversityRU2 papers
- Laboratoire d'Informatique, de Modélisation et d'Optimisation des SystèmesFR2 papers
- Pulkovo ObservatoryRU2 papers
- University of AmsterdamNL2 papers
- Accenture (Switzerland)CH1 paper
- Amsterdam University of the ArtsNL1 paper
- Astana Medical UniversityKZ1 paper
- Astro Space CenterRU1 paper
55 papers
Gated Bidirectional Linear Attention for Generative Retrieval
Artem Matveev, Vladislav Tytskiy, Sergei Makeev +1
In recommender systems, generative retrieval typically uses an encoder-decoder setup: an encoder processes a user interaction history, and an autoregressive decoder then generates…
fabric-lib: RDMA Point-to-Point Communication for LLM Systems
Nandor Licker, Kevin Hu, Vladimir Zaytsev +1
Emerging Large Language Model (LLM) system patterns, such as disaggregated inference, Mixture-of-Experts (MoE) routing, and asynchronous reinforcement fine-tuning, require flexible…
Blending Sequential Embeddings, Graphs, and Engineered Features: 4th Place Solution in RecSys Challenge 2025
Sergei Makeev, Alexandr Andreev, Vladimir Baikalov +3
This paper describes the 4th-place solution by team ambitious for the RecSys Challenge 2025, organized by Synerise and ACM RecSys, which focused on universal behavioral modeling. T…
eSASRec: Enhancing Transformer-based Recommendations in a Modular Fashion
Daria Tikhonovich, Nikita Zelinskiy, Aleksandr V. Petrov +4
Since their introduction, Transformer-based models, such as SASRec and BERT4Rec, have become common baselines for sequential recommendations, surpassing earlier neural and non-neur…
Correcting the LogQ Correction: Revisiting Sampled Softmax for Large-Scale Retrieval
Kirill Khrylchenko, Vladimir Baikalov, Sergei Makeev +2
Two-tower neural networks are a popular architecture for the retrieval stage in recommender systems. These models are typically trained with a softmax loss over the item catalog. H…
Scaling Recommender Transformers to One Billion Parameters
Kirill Khrylchenko, Artem Matveev, Sergei Makeev +1
While large transformer models have been successfully used in many real-world applications such as natural language processing, computer vision, and speech processing, scaling tran…