1 paper · 1 filter
Yang Xu, Chenang Li, Jiefu Zhang +3
Agent evaluations often use one benchmark to choose a workflow and then search for task types where its advantage weakens, so both conclusions are selected from the same data. We i…