1 paper
Siyuan Wang, Gaokai Zhang, Li Lyna Zhang +4
Reasoning over long contexts is essential for large language models. While reinforcement learning (RL) enhances short-context reasoning by inducing "Aha" moments in chain-of-though…