most citedYambda-5B -- A Large-Scale Multi-modal Dataset for Ranking And Retrieval

1 citations · 1 across the 4 of their papers we have counts for

collaborators

5 papers

cs.IR2026

Mitigating Collaborative Semantic ID Staleness in Generative Retrieval

Vladimir Baikalov, Iskander Bagautdinov, Sergey Muravyov

Generative retrieval with Semantic IDs (SIDs) assigns each item a discrete identifier and treats retrieval as a sequence generation problem rather than a nearest-neighbor search. W…

cs.IR2025

Blending Sequential Embeddings, Graphs, and Engineered Features: 4th Place Solution in RecSys Challenge 2025

Sergei Makeev, Alexandr Andreev, Vladimir Baikalov +3

This paper describes the 4th-place solution by team ambitious for the RecSys Challenge 2025, organized by Synerise and ACM RecSys, which focused on universal behavioral modeling. T…

cs.IR2025

Correcting the LogQ Correction: Revisiting Sampled Softmax for Large-Scale Retrieval

Kirill Khrylchenko, Vladimir Baikalov, Sergei Makeev +2

Two-tower neural networks are a popular architecture for the retrieval stage in recommender systems. These models are typically trained with a softmax loss over the item catalog. H…

cs.IR2025

Scaling Recommender Transformers to One Billion Parameters

Kirill Khrylchenko, Artem Matveev, Sergei Makeev +1

While large transformer models have been successfully used in many real-world applications such as natural language processing, computer vision, and speech processing, scaling tran…

cs.IR20251 cited

Yambda-5B -- A Large-Scale Multi-modal Dataset for Ranking And Retrieval

A. Ploshkin, V. Tytskiy, A. Pismenny +6

We present Yambda-5B, a large-scale open dataset sourced from the Yandex Music streaming platform. Yambda-5B contains 4.79 billion user-item interactions from 1 million users acros…