119 citations · 203 across the 6 of their papers we have counts for
7 papers
LoftQ: LoRA-Fine-Tuning-Aware Quantization for Large Language Models
Yixiao Li, Yifan Yu, Chen Liang +4
Quantization is an indispensable technique for serving Large Language Models (LLMs) and has recently found its way into LoRA fine-tuning. In this work we focus on the scenario wher…
Off-policy Confidence Sequences
Nikos Karampatziakis, Paul Mineiro, Aaditya Ramdas
We develop confidence bounds that hold uniformly over time for off-policy evaluation in the contextual bandit setting. These confidence sequences are based on recent ideas from mar…
Lessons from Contextual Bandit Learning in a Customer Support Bot
Nikos Karampatziakis, Sebastian Kochman, Jade Huang +3
In this work, we describe practical lessons we have learned from successfully using contextual bandits (CBs) to improve key business metrics of the Microsoft Virtual Agent for cust…
Empirical Likelihood for Contextual Bandits
Nikos Karampatziakis, John Langford, Paul Mineiro
We propose an estimator and confidence interval for computing the value of a policy from off-policy data in the contextual bandit setting. To this end we apply empirical likelihood…
Fast Label Embeddings for Extremely Large Output Spaces
Paul Mineiro, Nikos Karampatziakis
Many modern multiclass and multilabel problems are characterized by increasingly large output spaces. For these problems, label embeddings have been shown to be a useful primitive…
Efficient Optimal Learning for Contextual Bandits
Miroslav Dudik, Daniel Hsu, Satyen Kale +4
We address the problem of learning in an online setting where the learner repeatedly observes features, selects among a set of actions, and receives reward for the action taken. We…