2 papers
cs.CL2026
DARL: Encouraging Diverse Answers for General Reasoning without Verifiers
Chongxuan Huang, Lei Lin, Xiaodong Shi +2
Reinforcement Learning with Verifiable Rewards (RLVR) has demonstrated promising gains in enhancing the reasoning capabilities of large language models. However, its dependence on…
cs.CL2023
Layer-wise Representation Fusion for Compositional Generalization
Yafang Zheng, Lei Lin, Shuangtao Li +6
Existing neural models are demonstrated to struggle with compositional generalization (CG), i.e., the ability to systematically generalize to unseen compositions of seen components…