4 papers
DARL: Encouraging Diverse Answers for General Reasoning without Verifiers
Chongxuan Huang, Lei Lin, Xiaodong Shi +2
Reinforcement Learning with Verifiable Rewards (RLVR) has demonstrated promising gains in enhancing the reasoning capabilities of large language models. However, its dependence on…
From Tags to Trees: Structuring Fine-Grained Knowledge for Controllable Data Selection in LLM Instruction Tuning
Zihan Niu, Wenping Hu, Junmin Chen +3
Effective and controllable data selection is critical for LLM instruction tuning, especially with massive open-source datasets. Existing approaches primarily rely on instance-level…
SIN-Bench: Tracing Native Evidence Chains in Long-Context Multimodal Scientific Interleaved Literature
Yiming Ren, Junjie Wang, Yuxin Meng +11
Evaluating whether multimodal large language models truly understand long-form scientific papers remains challenging: answer-only metrics and synthetic "Needle-In-A-Haystack" tests…
MGFRec: Towards Reinforced Reasoning Recommendation with Multiple Groundings and Feedback
Shihao Cai, Chongming Gao, Haoyan Liu +4
The powerful reasoning and generative capabilities of large language models (LLMs) have inspired researchers to apply them to reasoning-based recommendation tasks, which require in…