2 papers
cs.AI2024
Towards Reliable Alignment: Uncertainty-aware RLHF
Debangshu Banerjee, Aditya Gopalan
Recent advances in aligning Large Language Models with human preferences have benefited from larger reward models and better preference data. However, most of these methodologies r…
cs.LG2023
On the Minimax Regret for Linear Bandits in a wide variety of Action Spaces
Debangshu Banerjee, Aditya Gopalan
As noted in the works of \cite{lattimore2020bandit}, it has been mentioned that it is an open problem to characterize the minimax regret of linear bandits in a wide variety of acti…