4 papers
FutureMind: Equipping Small Language Models with Strategic Thinking-Pattern Priors via Adaptive Knowledge Distillation
Shaoxiong Yang, Junting Li, Mengyuan Zhang +3
Small Language Models (SLMs) are attractive for cost-sensitive and resource-limited settings due to their efficient, low-latency inference. However, they often struggle with comple…
R^3: Replay, Reflection, and Ranking Rewards for LLM Reinforcement Learning
Zhizheng Jiang, Kang Zhao, Weikai Xu +5
Large reasoning models (LRMs) aim to solve diverse and complex problems through structured reasoning. Recent advances in group-based policy optimization methods have shown promise…
LightSearcher: Efficient DeepSearch via Experiential Memory
Hengzhi Lan, Yue Yu, Li Qian +5
DeepSearch paradigms have become a core enabler for deep reasoning models, allowing them to invoke external search tools to access up-to-date, domain-specific knowledge beyond para…
ICPO: Intrinsic Confidence-Driven Group Relative Preference Optimization for Efficient Reinforcement Learning
Jinpeng Wang, Chao Li, Ting Ye +3
Reinforcement Learning with Verifiable Rewards (RLVR) demonstrates significant potential in enhancing the reasoning capabilities of Large Language Models (LLMs). However, existing…