Showing cs.LGShow all
3 papers · 1 filter
cs.LG2026
POETS: Uncertainty-Aware LLM Optimization via Compute-Efficient Policy Ensembles
Nicolas Menet, Andreas Krause, Abbas Rahimi
Balancing exploration and exploitation is a core challenge in sequential decision-making and black-box optimization. We introduce POETS (licy nsembles for…
cs.LG2025
Thompson Sampling via Fine-Tuning of LLMs
Nicolas Menet, Aleksandar Terzić, Michael Hersche +2
Bayesian optimization in large unstructured discrete spaces is often hindered by the computational cost of maximizing acquisition functions due to the absence of gradients. We prop…
cs.LG2024
Bandits with Preference Feedback: A Stackelberg Game Perspective
Barna Pásztor, Parnian Kassraie, Andreas Krause
Bandits with preference feedback present a powerful tool for optimizing unknown target functions when only pairwise comparisons are allowed instead of direct value queries. This mo…