1 paper · 1 filter
Suho Shin, Chenghao Yang, Haifeng Xu +1
We introduce the tokenized linear bandit (TLB) and multi-armed bandit (TMAB), variants of linear and stochastic multi-armed bandit problems inspired by LLM decoding and alignment.…