1 paper · 1 filter
Dhruv Agarwal, Reece Adamson, Andrew McCallum +3
Open-ended scientific discovery with large language models (LLMs) increasingly operates as a long-horizon loop of hypothesis search and verification, where a reward signal guides w…