4 papers
Efficient Evaluation of LLM Performance with Statistical Guarantees
Skyler Wu, Yash Nair, Emmanuel J. Candès
Exhaustively evaluating many large language models (LLMs) on a large suite of benchmarks is expensive. We cast benchmarking as finite-population inference and, under a fixed query…
FUSE: Ensembling Verifiers with Zero Labeled Data
Joonhyuk Lee, Virginia Ma, Sarah Zhao +4
Verification of model outputs is rapidly emerging as a key primitive for both training and real-world deployment of large language models (LLMs). In practice, this often involves u…
Diversifying Conformal Selections
Yash Nair, Ying Jin, James Yang +1
When selecting from a list of potential candidates, it is important to ensure not only that the selected items are of high quality, but also that they are sufficiently dissimilar s…
Bernstein-von Mises for Adaptively Collected Data
Kevin Du, Yash Nair, Lucas Janson
Uncertainty quantification (UQ) for adaptively collected data, such as that coming from adaptive experiments, bandits, or reinforcement learning, is necessary for critical elements…