Showing cs.GTShow all
3 papers · 1 filter
cs.GT2026
Pluralistic Leaderboards
Nika Haghtalab, Ariel D. Procaccia, Han Shao +2
Recent leaderboard-based evaluations of large language models aggregate user feedback by fitting a Bradley--Terry model to pairwise comparisons, producing a single global ranking b…
cs.GT2025
The Disparate Effects of Partial Information in Bayesian Strategic Learning
Srikanth Avasarala, Serena Wang, Juba Ziani
We study how partial information about scoring rules affects fairness in strategic learning settings. In strategic learning, a learner deploys a scoring rule, and agents respond st…
cs.GT2024
Relying on the Metrics of Evaluated Agents
Serena Wang, Michael I. Jordan, Katrina Ligett +1
Online platforms and regulators face a continuing problem of designing effective evaluation metrics. While tools for collecting and processing data continue to progress, this has n…