activity
20242026
collaborators

5 papers

cs.LG2026

CollabEval: Statistically Efficient Collaborative Model Evaluation via Matrix Completion

Adam Fisch, Daniel Deutsch, Joshua Maynez +5

Evaluating generative AI models is a routine, but resource-intensive, process that is conducted over and over again during the course of model development. In this work, we propose…

math.ST2026

Multiple-Prediction-Powered Inference

Charlie Cowen-Breen, Alekh Agarwal, Stephen Bates +4

Statistical estimation often involves tradeoffs between expensive, high-quality measurements and a variety of lower-quality proxies. We introduce Multiple-Prediction-Powered Infere…

stat.ML2026

An Efficient Algorithm for Thresholding Monte Carlo Tree Search

Shoma Nameki, Atsuyoshi Nakamura, Junpei Komiyama +1

We introduce the Thresholding Monte Carlo Tree Search problem, in which, given a tree and a threshold , a player must answer whether the root node value of $\math…

cs.CL2025

The FACTS Leaderboard: A Comprehensive Benchmark for Large Language Model Factuality

Aileen Cheng, Alon Jacovi, Amir Globerson +62

We introduce The FACTS Leaderboard, an online leaderboard suite and associated set of benchmarks that comprehensively evaluates the ability of language models to generate factually…

cs.LG2024

Stratified Prediction-Powered Inference for Hybrid Language Model Evaluation

Adam Fisch, Joshua Maynez, R. Alex Hofer +3

Prediction-powered inference (PPI) is a method that improves statistical estimates based on limited human-labeled data. PPI achieves this by combining small amounts of human-labele…