1 paper · 1 filter
Yongcheng Zeng, Xinyu Cui, Yan Song +9
RLVR has substantially improved the reasoning capabilities of LLMs. However, existing methods typically parameterize temporal progression in the Markov Decision Process by token-by…