8 citations · 12 across the 7 of their papers we have counts for
7 papers
Provably Robust DPO: Aligning Language Models with Noisy Feedback
Sayak Ray Chowdhury, Anush Kini, Nagarajan Natarajan
Learning from preference-based feedback has recently gained traction as a promising approach to align language models with human interests. While these aligned generative models ha…
GAR-meets-RAG Paradigm for Zero-Shot Information Retrieval
Daman Arora, Anush Kini, Sayak Ray Chowdhury +3
Given a query and a document corpus, the information retrieval (IR) task is to output a ranked list of relevant documents. Combining large language models (LLMs) with embedding-bas…
Differentially Private Reward Estimation with Preference Feedback
Sayak Ray Chowdhury, Xingyu Zhou, Nagarajan Natarajan
Learning from preference-based feedback has recently gained considerable traction as a promising approach to align generative models with human interests. Instead of relying on num…
Differentially Private Episodic Reinforcement Learning with Heavy-tailed Rewards
Yulian Wu, Xingyu Zhou, Sayak Ray Chowdhury +1
In this paper, we study the problem of (finite horizon tabular) Markov decision processes (MDPs) with heavy-tailed rewards under the constraint of differential privacy (DP). Compar…
On Differentially Private Federated Linear Contextual Bandits
Xingyu Zhou, Sayak Ray Chowdhury
We consider cross-silo federated linear contextual bandit (LCB) problem under differential privacy, where multiple silos (agents) interact with the local users and communicate via…
Model Selection in Reinforcement Learning with General Function Approximations
Avishek Ghosh, Sayak Ray Chowdhury
We consider model selection for classic Reinforcement Learning (RL) environments -- Multi Armed Bandits (MABs) and Markov Decision Processes (MDPs) -- under general function approx…