activity
20242026
most citedSynthetic continued pretraining

1 citations · 1 across the 3 of their papers we have counts for

collaborators

7 papers

cs.CL2026

Towards Execution-Grounded Automated AI Research

Chenglei Si, Zitong Yang, Yejin Choi +3

Automated AI research holds great potential to accelerate scientific discovery. However, current LLMs often generate plausible-looking but ineffective ideas. Execution grounding ma…

stat.ML2025

Robust Sampling for Active Statistical Inference

Puheng Li, Tijana Zrnic, Emmanuel Candès

Active statistical inference is a new method for inference with AI-assisted data collection. Given a budget on the number of labeled data points that can be collected and assuming…

stat.ME2025

Imputation-Powered Inference

Sarah Zhao, Emmanuel Candès

Modern multi-modal and multi-site data frequently suffer from blockwise missingness, where subsets of features are missing for groups of individuals, creating complex patterns that…

cs.CL2025

Synthetic bootstrapped pretraining

Zitong Yang, Aonan Zhang, Hong Liu +4

We introduce Synthetic Bootstrapped Pretraining (SBP), a language model (LM) pretraining procedure that first learns a model of relations between documents from the pretraining dat…

stat.ML2025

Probably Approximately Correct Labels

Emmanuel J. Candès, Andrew Ilyas, Tijana Zrnic

Obtaining high-quality labeled datasets is often costly, requiring either human annotation or expensive experiments. In theory, powerful pre-trained AI models provide an opportunit…

cs.CL2025

s1: Simple test-time scaling

Niklas Muennighoff, Zitong Yang, Weijia Shi +7

Test-time scaling is a promising new approach to language modeling that uses extra test-time compute to improve performance. Recently, OpenAI's o1 model showed this capability but…