3 papers
cs.AI2025
What Does It Take to Be a Good AI Research Agent? Studying the Role of Ideation Diversity
Alexis Audran-Reiss, Jordi Armengol-Estapé, Karen Hambardzumyan +17
AI research agents offer the promise to accelerate scientific progress by automating the design, implementation, and training of machine learning models. However, the field is stil…
cs.LG2025
Eval Factsheets: A Structured Framework for Documenting AI Evaluations
Florian Bordes, Candace Ross, Justine T Kao +2
The rapid proliferation of benchmarks has created significant challenges in reproducibility, transparency, and informed decision-making. However, unlike datasets and models -- whic…
cs.CV2025
IntPhys 2: Benchmarking Intuitive Physics Understanding In Complex Synthetic Environments
Florian Bordes, Quentin Garrido, Justine T Kao +3
We present IntPhys 2, a video benchmark designed to evaluate the intuitive physics understanding of deep learning models. Building on the original IntPhys benchmark, IntPhys 2 focu…