1 paper · 1 filter
Moises Andrade, Joonhyuk Cha, Brandon Ho +3
Verifiers--functions assigning rewards to agent behavior--have been key to AI progress in math, code, and games. However, extending gains to domains without clear-cut success crite…