5 papers
CollabEval: Statistically Efficient Collaborative Model Evaluation via Matrix Completion
Adam Fisch, Daniel Deutsch, Joshua Maynez +5
Evaluating generative AI models is a routine, but resource-intensive, process that is conducted over and over again during the course of model development. In this work, we propose…
Multiple-Prediction-Powered Inference
Charlie Cowen-Breen, Alekh Agarwal, Stephen Bates +4
Statistical estimation often involves tradeoffs between expensive, high-quality measurements and a variety of lower-quality proxies. We introduce Multiple-Prediction-Powered Infere…
An Efficient Algorithm for Thresholding Monte Carlo Tree Search
Shoma Nameki, Atsuyoshi Nakamura, Junpei Komiyama +1
We introduce the Thresholding Monte Carlo Tree Search problem, in which, given a tree and a threshold , a player must answer whether the root node value of $\math…
The FACTS Leaderboard: A Comprehensive Benchmark for Large Language Model Factuality
Aileen Cheng, Alon Jacovi, Amir Globerson +62
We introduce The FACTS Leaderboard, an online leaderboard suite and associated set of benchmarks that comprehensively evaluates the ability of language models to generate factually…
Stratified Prediction-Powered Inference for Hybrid Language Model Evaluation
Adam Fisch, Joshua Maynez, R. Alex Hofer +3
Prediction-powered inference (PPI) is a method that improves statistical estimates based on limited human-labeled data. PPI achieves this by combining small amounts of human-labele…