4 citations · 6 across the 22 of their papers we have counts for
1 paper · 1 filter
Kailong Fan, Anqi Pu, Yichen Wu +7
Test-time reinforcement learning adapts a model on its own unlabeled test set using majority-vote pseudo-labels and has shown strong results in mathematics. We show that this recip…