2 citations · 2 across the 4 of their papers we have counts for
4 papers · 1 filter
PreScience: A Dataset and Benchmark for Scientific Forecasting
Anirudh Ajith, Amanpreet Singh, Jay DeYoung +7
Can AI systems trained on the existing scientific record forecast the advances that will follow? We introduce PreScience, a dataset and benchmark for scientific forecasting built a…
Because we have LLMs, we Can and Should Pursue Agentic Interpretability
Been Kim, John Hewitt, Neel Nanda +2
The era of Large Language Models (LLMs) presents a new opportunity for interpretability--agentic interpretability: a multi-turn conversation with an LLM wherein the LLM proactively…
CodeScientist: End-to-End Semi-Automated Scientific Discovery with Code-based Experimentation
Peter Jansen, Oyvind Tafjord, Marissa Radensky +6
Despite the surge of interest in autonomous scientific discovery (ASD) of software artifacts (e.g., improved ML algorithms), current ASD systems face two key limitations: (1) they…
DISCOVERYWORLD: A Virtual Environment for Developing and Evaluating Automated Scientific Discovery Agents
Peter Jansen, Marc-Alexandre Côté, Tushar Khot +5
Automated scientific discovery promises to accelerate progress across scientific domains. However, developing and evaluating an AI agent's capacity for end-to-end scientific reason…