2 papers
cs.CL2026
DARL: Encouraging Diverse Answers for General Reasoning without Verifiers
Chongxuan Huang, Lei Lin, Xiaodong Shi +2
Reinforcement Learning with Verifiable Rewards (RLVR) has demonstrated promising gains in enhancing the reasoning capabilities of large language models. However, its dependence on…
cs.CL2026
From Tags to Trees: Structuring Fine-Grained Knowledge for Controllable Data Selection in LLM Instruction Tuning
Zihan Niu, Wenping Hu, Junmin Chen +3
Effective and controllable data selection is critical for LLM instruction tuning, especially with massive open-source datasets. Existing approaches primarily rely on instance-level…