papers

Publications (26)

cs.AI2026

AstaBench: Rigorous Benchmarking of AI Agents with a Scientific Research Suite

Jonathan Bragg, Mike D'Arcy, Nishant Balepur +36

AI agents hold the potential to revolutionize scientific productivity by automating literature reviews, replicating experiments, analyzing data, and even proposing new directions o…

cs.CL2021

On the Challenges of Evaluating Compositional Explanations in Multi-Hop Inference: Relevance, Completeness, and Expert Ratings

Peter Jansen, Kelly Smith, Dan Moreno +1

Building compositional explanations requires models to combine two or more facts that, together, describe why the answer to a question is correct. Typically, these "multi-hop" expl…

cs.CL2023

From Words to Wires: Generating Functioning Electronic Devices from Natural Language Descriptions

Peter Jansen

In this work, we show that contemporary language models have a previously unknown skill -- the capacity for electronic circuit design from high-level textual descriptions, akin to…

cs.AI2025

Matter-of-Fact: A Benchmark for Verifying the Feasibility of Literature-Supported Claims in Materials Science

Peter Jansen, Samiah Hassan, Ruoyao Wang

Contemporary approaches to assisted scientific discovery use language models to automatically generate large numbers of potential hypothesis to test, while also automatically gener…

cs.CL2026

Generating Literature-Driven Scientific Theories at Scale

Peter Jansen, Peter Clark, Doug Downey +1

Contemporary automated scientific discovery has focused on agents for generating scientific experiments, while systems that perform higher-level scientific activities such as theor…

cs.CL2022

ScienceWorld: Is your Agent Smarter than a 5th Grader?

Ruoyao Wang, Peter Jansen, Marc-Alexandre Côté +1

We present ScienceWorld, a benchmark to test agents' scientific reasoning abilities in a new interactive text environment at the level of a standard elementary school science curri…