3 papers
cs.LG2025
Metritocracy: Representative Metrics for Lite Benchmarks
Ariel Procaccia, Benjamin Schiffer, Serena Wang +1
A common problem in LLM evaluation is how to choose a subset of metrics from a full suite of possible metrics. Subset selection is usually done for efficiency or interpretability r…
cs.LG2025
Clone-Robust AI Alignment
Ariel D. Procaccia, Benjamin Schiffer, Shirley Zhang
A key challenge in training Large Language Models (LLMs) is properly aligning them with human preferences. Reinforcement Learning with Human Feedback (RLHF) uses pairwise compariso…
cs.GT2025
Improved Regret Bounds for Online Fair Division with Bandit Learning
Benjamin Schiffer, Shirley Zhang
We study online fair division when there are a finite number of item types and the player values for the items are drawn randomly from distributions with unknown means. In this set…