2 papers
cs.AI2026
ReSyn: Autonomously Scaling Synthetic Environments for Reasoning Models
Andre He, Nathaniel Weir, Kaj Bostrom +4
Reinforcement learning with verifiable rewards (RLVR) has emerged as a promising approach for training reasoning language models (RLMs) by leveraging supervision from verifiers. Al…
cs.LG2025
Offline Learning and Forgetting for Reasoning with Large Language Models
Tianwei Ni, Allen Nie, Sapana Chaudhary +3
Leveraging inference-time search in large language models has proven effective in further enhancing a trained model's capability to solve complex mathematical and reasoning problem…