1 paper · 1 filter
Darius A. Faroughy, Sofia Palacios Schweitzer, Ian Pang +2
Autonomous language-model agents are increasingly evaluated on long-horizon tool-use tasks, but existing benchmarks rarely capture the complexity and nuance of real scientific work…