1 paper
Kunbin Xu, Xingzuo Li, Xuefeng Bai +1
Test-time reinforcement learning (TTRL) improves the reasoning capabilities of large language models without labeled data by updating the policy with pseudo-labels constructed thro…