From the 1 of 1 linked paper with an AI index.
1 paper
Bo-Wen Zhang, Junwei He, Wen Wang +5
The paper introduces CoRT, a method that uses counterfactual replay to assign token-level credit in rubric-guided reinforcement learning for language models, improving credit alloc…