2 papers
cs.LG2026
Learning What to Recommend: Minimax Optimal Simple Regret in Logistic Bandits
Shuai Liu, Alireza Bakhtiari, Alex Ayoub +2
We study stochastic logistic bandits with -dimensional action features under the simple-regret objective, where a learner uses rounds of exploration to output a single final…
cs.LG2024
Efficient Exploration for LLMs
Vikranth Dwaracherla, Seyed Mohammad Asghari, Botao Hao +1
We present evidence of substantial benefit from efficient exploration in gathering human feedback to improve large language models. In our experiments, an agent sequentially genera…