148 citations · 150 across the 9 of their papers we have counts for
1 paper · 2 filters
Ambuj Mehrish, Sebastiano Vascon
On unlabeled test data, reinforcement learning lacks a ground-truth reward; test-time RL methods derive one from the model's own roll-outs, rewarding those that match the majority…