activity
20242026
collaborators

9 papers

cs.LG2026

Strategic Feature Selection

Jivat Neet Kaur, Pratik Patil, Divya Shanmugam +6

When algorithmic predictors inform resource allocation in high-stakes domains such as healthcare, these predictors must account for strategic manipulation of input features. The ty…

cs.GT2026

Pluralistic Leaderboards

Nika Haghtalab, Ariel D. Procaccia, Han Shao +2

Recent leaderboard-based evaluations of large language models aggregate user feedback by fitting a Bradley--Terry model to pairwise comparisons, producing a single global ranking b…

cs.LG2026

Robust AI Evaluation through Maximal Lotteries

Hadi Khalaf, Serena L. Wang, Daniel Halpern +3

The standard way to evaluate language models on subjective tasks is through pairwise comparisons: an annotator chooses the "better" of two responses to a prompt. Leaderboards aggre…

cs.SE2025

Incentives and Outcomes in Bug Bounties

Serena Wang, Martino Banchio, Krzysztof Kotowicz +3

Bug bounty programs have contributed significantly to security in technology firms in the last decade, but little is known about the role of reward incentives in producing useful o…

cs.LG2025

Metritocracy: Representative Metrics for Lite Benchmarks

Ariel Procaccia, Benjamin Schiffer, Serena Wang +1

A common problem in LLM evaluation is how to choose a subset of metrics from a full suite of possible metrics. Subset selection is usually done for efficiency or interpretability r…

cs.GT2025

The Disparate Effects of Partial Information in Bayesian Strategic Learning

Srikanth Avasarala, Serena Wang, Juba Ziani

We study how partial information about scoring rules affects fairness in strategic learning settings. In strategic learning, a learner deploys a scoring rule, and agents respond st…