4 papers · 1 filter
Cost Efficient Fairness Audit Under Partial Feedback
Nirjhar Das, Mohit Sharma, Praharsh Nanavati +2
We study the problem of auditing the fairness of a given classifier under partial feedback, where true labels are available only for positively classified individuals, (e.g., loan…
Generalized Linear Bandits with Limited Adaptivity
Ayush Sawarni, Nirjhar Das, Siddharth Barman +1
We study the generalized linear contextual bandit problem within the constraints of limited adaptivity. In this paper, we present two algorithms, and $\texttt{R…
Active Preference Optimization for Sample Efficient RLHF
Nirjhar Das, Souradip Chakraborty, Aldo Pacchiano +1
Large Language Models (LLMs) aligned using Reinforcement Learning from Human Feedback (RLHF) have shown remarkable generation abilities in numerous tasks. However, collecting high-…
Linear Contextual Bandits with Hybrid Payoff: Revisited
Nirjhar Das, Gaurav Sinha
We study the Linear Contextual Bandit problem in the hybrid reward setting. In this setting every arm's reward model contains arm specific parameters in addition to parameters shar…