2 papers
cs.LG2026
Vector Policy Optimization: Training for Diversity Improves Test-Time Search
Ryan Bahlous-Boldi, Isha Puri, Idan Shenfeld +6
Language models must now generalize out of the box to novel environments and work inside inference-scaling search procedures, such as AlphaEvolve, that select rollouts with a varie…
cs.LG2025
Going Beyond Heuristics by Imposing Policy Improvement as a Constraint
Chi-Chang Lee, Zhang-Wei Hong, Pulkit Agrawal
In many reinforcement learning (RL) applications, augmenting the task rewards with heuristic rewards that encode human priors about how a task should be solved is crucial for achie…