1 paper · 1 filter
Hantao Yang, Hong Xie, Defu Lian +1
This paper revisits the LLM cache bandit problem, with a special focus on addressing the query heterogeneity for cost-effective LLM inference. Previous works often assume uniform q…