From the 1 of 1 linked paper with an AI index.
1 paper
Qiangqiang He, Zhongheng Wu, ZiJian Wang
The paper examines how uniformly assigning credit to all tokens during reinforcement learning for long chain‑of‑thought reasoning can be misleading, and introduces Counterfactual S…