2 papers
cs.LG2026
Where Hindsight Credit Can Reside: A Signed-Capacity View of Token Updates in RLVR
Yuhang He, Haodong Wu, Siyi Liu +7
Reinforcement Learning with Verifiable Rewards (RLVR) improves the reasoning ability of Large Language Models (LLMs), but sparse outcome rewards make token-level credit assignment…
cs.CL2024
Overview of AI-Debater 2023: The Challenges of Argument Generation Tasks
Jiayu Lin, Guanrong Chen, Bojun Jin +26
In this paper we present the results of the AI-Debater 2023 Challenge held by the Chinese Conference on Affect Computing (CCAC 2023), and introduce the related datasets. We organiz…