3 papers
cs.LG2026
Not All Negative Samples Are Equal: LLMs Learn Better from Plausible Reasoning
Zixiang Di, Jinyi Han, Shuo Zhang +8
Learning from negative samples holds great promise for improving Large Language Model (LLM) reasoning capability, yet existing methods treat all incorrect responses as equally info…
cs.LG2026
From Atoms to Chains: Divergence-Guided Reasoning Curriculum for Unlabeled LLM Domain Adaptation
Yongqi Wang, Xiaofeng Ji, Jie Wang +6
Adapting Large Language Models (LLMs) to specialized domains without human-annotated data is a crucial yet formidable challenge. Widely adopted knowledge distillation methods often…
cs.LG2025
CURE: Critical-Token-Guided Re-Concatenation for Entropy-Collapse Prevention
Qingbin Li, Rongkun Xue, Jie Wang +8
Recent advances in Reinforcement Learning with Verified Reward (RLVR) have driven the emergence of more sophisticated cognitive behaviors in large language models (LLMs), thereby e…