5 papers
As It Was: Aligning LLM Search Evaluation with Historical User Preferences
Ali Vardasbi, Gustavo Penha, Enrico Palumbo +3
Large-scale search systems evolve faster than human quality assurance can scale, especially for long-tail intents and multilingual queries. LLM-as-a-judge approaches provide a scal…
From IR to RecSys: Evaluating LLM-based Judges in Cranfield-style Recommendation Collections
Gustavo Penha, Aleksandr V. Petrov, Claudia Hauff +9
The Cranfield paradigm has long provided reliable, reproducible evaluation in ad hoc retrieval, and recent work has begun extending this framework to recommender systems. A recent…
Semantic IDs for Joint Generative Search and Recommendation
Gustavo Penha, Edoardo D'Amico, Marco De Nadai +8
Generative models powered by Large Language Models (LLMs) are emerging as a unified solution for powering both recommendation and search tasks. A key design choice in these models…
Adaptive Repetition for Mitigating Position Bias in LLM-Based Ranking
Ali Vardasbi, Gustavo Penha, Claudia Hauff +1
When using LLMs to rank items based on given criteria, or evaluate answers, the order of candidate items can influence the model's final decision. This sensitivity to item position…
Bridging Search and Recommendation in Generative Retrieval: Does One Task Help the Other?
Gustavo Penha, Ali Vardasbi, Enrico Palumbo +2
Generative retrieval for search and recommendation is a promising paradigm for retrieving items, offering an alternative to traditional methods that depend on external indexes and…