Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
Interactive Multi-Objective Probabilistic Preference Learning with Soft and Hard Bounds
Edward Chen, Sang T. Truong, Natalie Dullerud +2
High-stakes decision-making involves navigating multiple competing objectives with expensive evaluations. For instance, in brachytherapy, clinicians must balance maximizing tumor c…
cs.AI2026
Metric Match: A Subset Selection Approach to Evaluating LLM Judge Reliability
Alyssa Unell, Natalie Dullerud, Naomi Boneh +4
LLM judges are used to reduce the need for costly human labor in evaluating open-ended text generation. However, the reliability of these judges depends critically on their alignme…