activity
20242026
collaborators

6 papers

cs.AI2026

What Fits (Into Few Tokens) Doesn't Overfit: Compression and Generalization in ML Research Agents

Martin Andres Bertran, Aaron Roth, Zhiwei Steven Wu

Reusing a held-out benchmark adaptively should, in principle, invite overfitting. Yet benchmark-driven machine learning (ML) has produced surprisingly little overfitting in practic…

cs.AI2026

Many AI Analysts, One Dataset: Navigating the Agentic Data Science Multiverse

Martin Bertran, Riccardo Fogliato, Zhiwei Steven Wu

Empirical conclusions depend not only on data but on analytic decisions made throughout the research process. Many-analyst studies have quantified this dependence: independent team…

stat.ME2025

Stronger Neyman Regret Guarantees for Adaptive Experimental Design

Georgy Noarov, Riccardo Fogliato, Martin Bertran +1

We study the design of adaptive, sequential experiments for unbiased average treatment effect (ATE) estimation in the design-based potential outcomes setting. Our goal is to develo…

cs.LG2024

Improving LLM Group Fairness on Tabular Data via In-Context Learning

Valeriia Cherepanova, Chia-Jung Lee, Nil-Jana Akpinar +4

Large language models (LLMs) have been shown to be effective on tabular prediction tasks in the low-data regime, leveraging their internal knowledge and ability to learn from instr…

cs.LG2024

Order of Magnitude Speedups for LLM Membership Inference

Rongting Zhang, Martin Bertran, Aaron Roth

Large Language Models (LLMs) have the promise to revolutionize computing broadly, but their complexity and extensive training data also expose significant privacy vulnerabilities.…

cs.LG2024

Reconstruction Attacks on Machine Unlearning: Simple Models are Vulnerable

Martin Bertran, Shuai Tang, Michael Kearns +3

Machine unlearning is motivated by desire for data autonomy: a person can request to have their data's influence removed from deployed models, and those models should be updated as…