2 papers
cs.LG2026
Policy Regret for Embedding Model Routing: Contextual Bandits with Low-Rank Experts
Yan Dai, Negin Golrezaei, Patrick Jaillet
Modern recommendation systems increasingly rely on dynamically routing diverse queries to multiple embedding models. Despite its practical significance, this problem remains poorly…
cs.LG2025
uniINF: Best-of-Both-Worlds Algorithm for Parameter-Free Heavy-Tailed MABs
Yu Chen, Jiatai Huang, Yan Dai +1
In this paper, we present a novel algorithm, uniINF, for the Heavy-Tailed Multi-Armed Bandits (HTMAB) problem, demonstrating robustness and adaptability in both stochastic and adve…