From the 1 of 1 linked paper with an AI index.
1 paper
Jianze Wang, Kunwang Zheng, Ying Liu +5
The paper introduces SERPO, a test-time reinforcement learning approach that lets language models self‑improve during inference by jointly evolving response evidence, query‑specifi…