3 papers
cs.CL2026
BiPACE: Bisimulation-Guided Policy Optimization with Action Counterfactual Estimation for LLM Agents
Hanyang Wang, Weijieying Ren, Yuxiang Zhang +4
Stepwise group-based RL is an attractive way to train long-horizon LLM agents without a learned critic: it reuses multiple sampled rollouts to estimate local advantages. Its weakne…
cs.LG2026
Bayesian Optimization with Preference Exploration using a Monotonic Neural Network Ensemble
Hanyang Wang, Juergen Branke, Matthias Poloczek
Many real-world black-box optimization problems have multiple conflicting objectives. Rather than attempting to approximate the entire set of Pareto-optimal solutions, interactive…
cs.LG2025
Respecting the limit:Bayesian optimization with a bound on the optimal value
Hanyang Wang, Juergen Branke, Matthias Poloczek
In many real-world optimization problems, we have prior information about what objective function values are achievable. In this paper, we study the scenario that we have either ex…