1 paper
Grégoire Martinon, Ibrahim Merad, Mohammed Raki
Reliable evaluation of agentic systems requires unbiased estimates with valid uncertainty, but standard practice navigates between costly human annotation and biased LLM-as-judge p…