1 paper · 1 filter
Ambuj Mehrish, Sebastiano Vascon
On unlabeled test data, reinforcement learning lacks a ground-truth reward; test-time RL methods derive one from the model's own roll-outs, rewarding those that match the majority…