1 paper
Jingjie Ning, Xueqi Li, Yibo Kong +1
AI research agents need reliable knowledge of how their experiments change outcomes. We introduce WhatWorkedBench to measure experimental understanding, the accuracy of predictions…