Showing cs.AIShow all
3 papers · 1 filter
cs.AI2026
EvoThink: Evolving Thinking in Large Reasoning Models via Self-Pruning and Aha-Moment Preference Optimization
Xinbang Dai, Zheyu Xin, Huikang Hu +7
Large Reasoning Models (LRMs) often suffer from overthinking due to redundant verification steps. Existing approaches for mitigating overthinking, such as fast-slow thinking switch…
cs.AI2025
Reinforcement Learning Foundations for Deep Research Systems: A Survey
Wenjun Li, Zhi Chen, Jingru Lin +8
Deep research systems, agentic AI that solve complex, multi-step tasks by coordinating reasoning, search across the open web and user files, and tool use, are moving toward hierarc…
cs.AI2024
Aligning Crowd Feedback via Distributional Preference Reward Modeling
Dexun Li, Cong Zhang, Kuicai Dong +3
Deep Reinforcement Learning is widely used for aligning Large Language Models (LLM) with human preference. However, the conventional reward modelling is predominantly dependent on…