Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
What Fits (Into Few Tokens) Doesn't Overfit: Compression and Generalization in ML Research Agents
Martin Andres Bertran, Aaron Roth, Zhiwei Steven Wu
Reusing a held-out benchmark adaptively should, in principle, invite overfitting. Yet benchmark-driven machine learning (ML) has produced surprisingly little overfitting in practic…
cs.AI2026
Many AI Analysts, One Dataset: Navigating the Agentic Data Science Multiverse
Martin Bertran, Riccardo Fogliato, Zhiwei Steven Wu
Empirical conclusions depend not only on data but on analytic decisions made throughout the research process. Many-analyst studies have quantified this dependence: independent team…