1 paper · 1 filter
Wei Xiong, Hanning Zhang, Chenlu Ye +3
We study self-rewarding reasoning large language models (LLMs), which can simultaneously generate step-by-step reasoning and evaluate the correctness of their outputs during the in…