4 papers
Off-Policy Evaluation from Logged Human Feedback
Aniruddha Bhargava, Lalit Jain, Branislav Kveton +2
Learning from human feedback has been central to recent advances in artificial intelligence and machine learning. Since the collection of human feedback is costly, a natural questi…
Experimental Design for Active Transductive Inference in Large Language Models
Subhojyoti Mukherjee, Anusha Lalitha, Aniket Deshmukh +3
One emergent ability of large language models (LLMs) is that query-specific examples can be included in the prompt at inference time. In this work, we use active learning for adapt…
Optimal Design for Human Preference Elicitation
Subhojyoti Mukherjee, Anusha Lalitha, Kousha Kalantari +4
Learning of preference models from human feedback has been central to recent advances in artificial intelligence. Motivated by the cost of obtaining high-quality human annotations,…
Pessimistic Off-Policy Multi-Objective Optimization
Shima Alizadeh, Aniruddha Bhargava, Karthick Gopalswamy +3
Multi-objective optimization is a type of decision making problems where multiple conflicting objectives are optimized. We study offline optimization of multi-objective policies fr…