Showing cs.LGShow all
3 papers · 1 filter
cs.LG2026
Strategic Feature Selection
Jivat Neet Kaur, Pratik Patil, Divya Shanmugam +6
When algorithmic predictors inform resource allocation in high-stakes domains such as healthcare, these predictors must account for strategic manipulation of input features. The ty…
cs.LG2026
Robust AI Evaluation through Maximal Lotteries
Hadi Khalaf, Serena L. Wang, Daniel Halpern +3
The standard way to evaluate language models on subjective tasks is through pairwise comparisons: an annotator chooses the "better" of two responses to a prompt. Leaderboards aggre…
cs.LG2025
Metritocracy: Representative Metrics for Lite Benchmarks
Ariel Procaccia, Benjamin Schiffer, Serena Wang +1
A common problem in LLM evaluation is how to choose a subset of metrics from a full suite of possible metrics. Subset selection is usually done for efficiency or interpretability r…