11 papers
As It Was: Aligning LLM Search Evaluation with Historical User Preferences
Ali Vardasbi, Gustavo Penha, Enrico Palumbo +3
Large-scale search systems evolve faster than human quality assurance can scale, especially for long-tail intents and multilingual queries. LLM-as-a-judge approaches provide a scal…
Deploying Semantic ID-based Generative Retrieval for Large-Scale Podcast Discovery at Spotify
Edoardo D'Amico, Marco De Nadai, Praveen Chandar +41
Podcast listening is often grounded in a set of favorite shows, while listener intent can evolve over time. This combination of stable preferences and changing intent motivates rec…
Stream: Scaling up Mechanistic Interpretability to Long Context in LLMs via Sparse Attention
J Rosser, José Luis Redondo GarcÃa, Gustavo Penha +2
As Large Language Models (LLMs) scale to million-token contexts, traditional Mechanistic Interpretability techniques for analyzing attention scale quadratically with context length…
From IR to RecSys: Evaluating LLM-based Judges in Cranfield-style Recommendation Collections
Gustavo Penha, Aleksandr V. Petrov, Claudia Hauff +9
The Cranfield paradigm has long provided reliable, reproducible evaluation in ad hoc retrieval, and recent work has begun extending this framework to recommender systems. A recent…
AudioBoost: Increasing Audiobook Retrievability in Spotify Search with Synthetic Query Generation
Enrico Palumbo, Gustavo Penha, Alva Liu +6
Spotify has recently introduced audiobooks as part of its catalog, complementing its music and podcast offering. Search is often the first entry point for users to access new items…
Semantic IDs for Joint Generative Search and Recommendation
Gustavo Penha, Edoardo D'Amico, Marco De Nadai +8
Generative models powered by Large Language Models (LLMs) are emerging as a unified solution for powering both recommendation and search tasks. A key design choice in these models…