1 paper · 1 filter
Lanbo Lin, Jiayao Liu, Tianyuan Yang +5
Evaluating agentic AI on open-ended professional tasks faces a fundamental dilemma between rigor and flexibility. Static rubrics provide rigorous, reproducible assessment but fail…