1 paper
Huchen Jiang, Yangyang Ma, Chaofan Ding +2
With current state-of-the-art approaches aimed at enhancing the reasoning capabilities of Large Language Models(LLMs) through iterative preference learning inspired by AlphaZero, w…