Showing cs.AIShow all
3 papers · 1 filter
cs.AI2026
LiteResearcher: A Scalable Agentic RL Training Framework for Deep Research Agent
Wanli Li, Bince Qu, Bo Pan +5
Reinforcement Learning (RL) has emerged as a powerful training paradigm for LLM-based agents. However, scaling agentic RL for deep research remains constrained by two coupled chall…
cs.AI2026
Emergent Slow Thinking in LLMs as Inverse Tree Freezing
Sihan Hu, Xiansheng Cai, Yuan Huang +5
Reinforcement learning with verifiable rewards (RLVR) enables large language models to acquire slow, multi-step reasoning from sparse final-answer signals. We provide a statistical…
cs.AI2026
Inverse Knowledge Search over Verifiable Reasoning: Synthesizing a Scientific Encyclopedia from a Long Chains-of-Thought Knowledge Base
Yu Li, Yuan Huang, Tao Wang +19
Most scientific materials compress reasoning, presenting conclusions while omitting the derivational chains that justify them. This compression hinders verification by lacking expl…