Showing cs.LGShow all
2 papers · 1 filter
cs.LG2025
MiGrATe: Mixed-Policy GRPO for Adaptation at Test-Time
Peter Phan, Dhruv Agarwal, Kavitha Srinivas +3
Large language models (LLMs) are increasingly being applied to black-box optimization tasks, from program synthesis to molecule design. Prior work typically leverages in-context le…
cs.LG2025
AutoDiscovery: Open-ended Scientific Discovery via Bayesian Surprise
Dhruv Agarwal, Bodhisattwa Prasad Majumder, Reece Adamson +8
The promise of autonomous scientific discovery (ASD) hinges not only on answering questions, but also on knowing which questions to ask. Most recent works in ASD explore the use of…