5 papers
As It Was: Aligning LLM Search Evaluation with Historical User Preferences
Ali Vardasbi, Gustavo Penha, Enrico Palumbo +3
Large-scale search systems evolve faster than human quality assurance can scale, especially for long-tail intents and multilingual queries. LLM-as-a-judge approaches provide a scal…
From IR to RecSys: Evaluating LLM-based Judges in Cranfield-style Recommendation Collections
Gustavo Penha, Aleksandr V. Petrov, Claudia Hauff +9
The Cranfield paradigm has long provided reliable, reproducible evaluation in ad hoc retrieval, and recent work has begun extending this framework to recommender systems. A recent…
Adaptive Repetition for Mitigating Position Bias in LLM-Based Ranking
Ali Vardasbi, Gustavo Penha, Claudia Hauff +1
When using LLMs to rank items based on given criteria, or evaluate answers, the order of candidate items can influence the model's final decision. This sensitivity to item position…
Policy-as-Prompt: Rethinking Content Moderation in the Age of Large Language Models
Konstantina Palla, José Luis Redondo GarcÃa, Claudia Hauff +5
Content moderation plays a critical role in shaping safe and inclusive online environments, balancing platform standards, user expectations, and regulatory frameworks. Traditionall…
PODTILE: Facilitating Podcast Episode Browsing with Auto-generated Chapters
Azin Ghazimatin, Ekaterina Garmash, Gustavo Penha +14
Listeners of long-form talk-audio content, such as podcast episodes, often find it challenging to understand the overall structure and locate relevant sections. A practical solutio…