34 citations · 43 across the 9 of their papers we have counts for
10 papers
Last Translation Benchmark
Vilém Zouhar, Niyati Bafna, Mukund Choudhary +241
For scientific progress, we need benchmarks that test the limits of state-of-the-art models, and evaluation methods that inform us about failure cases. As models get stronger, stan…
LegalPincite: Multi-level Legal Information Retrieval Dataset
Theresia Veronika Rampisela, Henrik Palmer Olsen, Giovanni Colavizza
A common task in legal Information Retrieval (IR) is to find relevant legal sources from case-law collections. While legal practice often requires pinpoint citations (pincites) to…
Offline Evaluation Measures of Fairness in Recommender Systems
Theresia Veronika Rampisela
The evaluation of recommender system fairness has become increasingly important, especially with recent legislation that emphasises the development of fair and responsible artifici…
Can Fairness Be Prompted? Prompt-Based Debiasing Strategies in High-Stakes Recommendations
Mihaela Rotar, Theresia Veronika Rampisela, Maria Maistro
Large Language Models (LLMs) can infer sensitive attributes such as gender or age from indirect cues like names and pronouns, potentially biasing recommendations. While several deb…
Measuring Individual User Fairness with User Similarity and Effectiveness Disparity
Theresia Veronika Rampisela, Maria Maistro, Tuukka Ruotsalo +1
Individual user fairness is commonly understood as treating similar users similarly. In Recommender Systems (RSs), several evaluation measures exist for quantifying individual user…
The Quest for Reliable Metrics of Responsible AI
Theresia Veronika Rampisela, Maria Maistro, Tuukka Ruotsalo +1
The development of Artificial Intelligence (AI), including AI in Science (AIS), should be done following the principles of responsible AI. Progress in responsible AI is often quant…