12 citations · 21 across the 10 of their papers we have counts for
6 papers · 1 filter
Robust Sampling for Active Statistical Inference
Puheng Li, Tijana Zrnic, Emmanuel Candès
Active statistical inference is a new method for inference with AI-assisted data collection. Given a budget on the number of labeled data points that can be collected and assuming…
Imputation-Powered Inference
Sarah Zhao, Emmanuel Candès
Modern multi-modal and multi-site data frequently suffer from blockwise missingness, where subsets of features are missing for groups of individuals, creating complex patterns that…
Synthetic bootstrapped pretraining
Zitong Yang, Aonan Zhang, Hong Liu +4
We introduce Synthetic Bootstrapped Pretraining (SBP), a language model (LM) pretraining procedure that first learns a model of relations between documents from the pretraining dat…
The Future of Artificial Intelligence and the Mathematical and Physical Sciences (AI+MPS)
Andrew Ferguson, Marisa LaFleur, Lars Ruthotto +97
This community paper developed out of the NSF Workshop on the Future of Artificial Intelligence (AI) and the Mathematical and Physics Sciences (MPS), which was held in March 2025 w…
Probably Approximately Correct Labels
Emmanuel J. Candès, Andrew Ilyas, Tijana Zrnic
Obtaining high-quality labeled datasets is often costly, requiring either human annotation or expensive experiments. In theory, powerful pre-trained AI models provide an opportunit…
s1: Simple test-time scaling
Niklas Muennighoff, Zitong Yang, Weijia Shi +7
Test-time scaling is a promising new approach to language modeling that uses extra test-time compute to improve performance. Recently, OpenAI's o1 model showed this capability but…