1 citations · 1 across the 3 of their papers we have counts for
1 paper · 1 filter
Aayam Bansal, Keertan Balaji
Existing benchmarks for autonomous AI scientists evaluate only final outputs---generated code, hypotheses, or papers---yet discard the reasoning process by which those outputs were…