paper

Online Learning with LLM Experts from Limited Feedback

arXiv:2609.05820

Abstract

We study adaptive routing of prompts to large language model (LLM) experts to maximize response quality in an online setting with limited feedback. We formulate it as a bandit problem with actions that represent experts and features that encode prompts, over a horizon of rounds. We propose algorithms that strategically select and observe rewards to minimize regret. In the full-information setting, we achieve a regret of , while in the bandit setting we achieve , where is a budget on feedback. Our experiments show that we efficiently learn high-quality routing strategies across diverse LLMs from limited feedback.

21 pages, 4 figures