6 papers
Robust LLM Performance Certification via Constrained Maximum Likelihood Estimation
Minghe Shen, Ananth Balashankar, Adam Fisch +2
The ability to rigorously estimate the failure rates of large language models (LLMs) is a prerequisite for their safe deployment. Currently, however, practitioners often face a tra…
Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
Gheorghe Comanici, Eric Bieber, Mike Schaekermann +3431
In this report, we introduce the Gemini 2.X model family: Gemini 2.5 Pro and Gemini 2.5 Flash, as well as our earlier Gemini 2.0 Flash and Flash-Lite models. Gemini 2.5 Pro is our…
Understanding challenges to the interpretation of disaggregated evaluations of algorithmic fairness
Stephen R. Pfohl, Natalie Harris, Chirag Nagpal +12
Disaggregated evaluation across subgroups is critical for assessing the fairness of machine learning models, but its uncritical use can mislead practitioners. We show that equal pe…
Prompts Generalize with Low Data: Non-vacuous Generalization Bounds for Optimizing Prompts with More Informative Priors
David Madras, Joshua Safyan, Qiuyi +1
Many prompt engineering techniques have been successful in practice, even when optimizing over a large prompt space with with a small amount of task-specific data. Recent work has…
Regression for the Mean: Auto-Evaluation and Inference with Few Labels through Post-hoc Regression
Benjamin Eyre, David Madras
The availability of machine learning systems that can effectively perform arbitrary tasks has led to synthetic labels from these systems being used in applications of statistical i…
QuEst: Enhancing Estimates of Quantile-Based Distributional Measures Using Model Predictions
Zhun Deng, Thomas P Zollo, Benjamin Eyre +3
As machine learning models grow increasingly competent, their predictions can supplement scarce or expensive data in various important domains. In support of this paradigm, algorit…