6 citations · 18 across the 9 of their papers we have counts for
1 paper · 1 filter
Shubham Gandhi, Saurabh Goyal, Kiran Kate +1
Reinforcement Learning from Verifiable Rewards works well when a task has a programmatic checker, but most long-horizon agent domains have none. We work in the outcome-blind settin…