2 papers
cs.AI2026
DeepLook: Deeper Thinking with Lookahead
Tingxin Yang, Zefeng Wang, Mengyue Wang +2
Inference-time scaling has emerged as a powerful paradigm for improving large language model reasoning, often delivering larger gains on difficult reasoning tasks than parameter sc…
cs.LG2026
EchoRL: Reinforcement Learning via Rollout Echoing
Jinhe Bi, Aniri, Minglai Yang +9
Reinforcement Learning with Verifiable Rewards is an effective route for post-training to strengthen the reasoning capability of large language models. However, as training proceed…