Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
Trading Human Curation for Synthetic Augmentation in RLVR
Akshansh, Leonardo Rosa Rodrigues, Michael Korostelev +2
The supply of high-quality training tasks is a central bottleneck for reinforcement learning from verifiable rewards (RLVR) on agentic language models. Each task requires a sandbox…
cs.LG2026
LEAP: Trajectory-Level Evaluation of LLMs in Iterative Scientific Design
Marilyn Zhang, Tianfeng Chen, Fabián Barzuna +2
LLMs are increasingly deployed in autonomous laboratories, under the assumption that their domain priors and reasoning over iterative feedback let them converge on good designs in…