1 paper · 1 filter
Renos Zabounidis, Aditya Golatkar, Michael Kleinman +3
We propose Re-FORC, an adaptive reward prediction method that, given a query, enables prediction of the expected future rewards as a function of the number of future thinking token…