1 paper · 1 filter
Nicolas Emmenegger, Ellery Stahler, Chara Podimata
Many applications require statistically valid inference across many related tasks, while using only a handful of high-quality labels per hypothesis. In AI evaluation, these tasks m…