1 paper
Weiqin Wang, Yile Wang, Kehao Chen +1
Test-time reinforcement learning mitigates the reliance on annotated data by using majority voting results as pseudo-labels, emerging as a complementary direction to reinforcement…