works on

From the 1 of 6 linked papers with an AI index.

most citedScaling Recommender Transformers to One Billion Parameters

1 citations · 1 across the 4 of their papers we have counts for

collaborators

6 papers

cs.IR2026

Embedding Items at Scale: Comparing GNN-Based and ID-Based Item Embeddings in the Yandex Ecosystem

Sergei Makeev, Artem Matveev, Vladimir Baikalov +1

The paper compares pretrained graph neural network item embeddings with end‑to‑end trainable embeddings in transformer‑based sequential recommendation systems at Yandex, finding pr…

cs.IR2026

Session-Level Optimization for Large-Scale Retrieval using REINFORCE with Multi-Step Off-Policy Correction

Artem Matveev, Sergei Makeev, Aleksei Krasilnikov +3

Two-tower models are a widely used paradigm for large-scale retrieval in recommendation. However, they are typically trained with myopic supervised objectives, such as next-item pr…

cs.IR2026

Variable-Length Semantic IDs for Recommender Systems

Kirill Khrylchenko

Generative models are increasingly used in recommender systems, both for modeling user behavior as event sequences and for integrating large language models into recommendation pip…

cs.IR20261 cited

Scaling Recommender Transformers to One Billion Parameters

Kirill Khrylchenko, Artem Matveev, Sergei Makeev +1

While large transformer models have been successfully used in many real-world applications such as natural language processing, computer vision, and speech processing, scaling tran…

cs.IR2025

Blending Sequential Embeddings, Graphs, and Engineered Features: 4th Place Solution in RecSys Challenge 2025

Sergei Makeev, Alexandr Andreev, Vladimir Baikalov +3

This paper describes the 4th-place solution by team ambitious for the RecSys Challenge 2025, organized by Synerise and ACM RecSys, which focused on universal behavioral modeling. T…

cs.IR2025

Correcting the LogQ Correction: Revisiting Sampled Softmax for Large-Scale Retrieval

Kirill Khrylchenko, Vladimir Baikalov, Sergei Makeev +2

Two-tower neural networks are a popular architecture for the retrieval stage in recommender systems. These models are typically trained with a softmax loss over the item catalog. H…