Showing cs.LGShow all
3 papers · 1 filter
cs.LG2026
POETS: Uncertainty-Aware LLM Optimization via Compute-Efficient Policy Ensembles
Nicolas Menet, Andreas Krause, Abbas Rahimi
Balancing exploration and exploitation is a core challenge in sequential decision-making and black-box optimization. We introduce POETS (licy nsembles for…
cs.LG2026
Thompson Sampling via Fine-Tuning of LLMs
Nicolas Menet, Aleksandar TerziÄ, Michael Hersche +2
Bayesian optimization in large unstructured discrete spaces is often hindered by the computational cost of maximizing acquisition functions due to the absence of gradients. We prop…
cs.LG2025
Bandits with Preference Feedback: A Stackelberg Game Perspective
Barna Pásztor, Parnian Kassraie, Andreas Krause
Bandits with preference feedback present a powerful tool for optimizing unknown target functions when only pairwise comparisons are allowed instead of direct value queries. This mo…