10 papers
MINT: A Universal Zero-Shot Predictor for Transaction Data
Parameswaran Kamalaruban, Viktor Drobnyi, Maeve Madigan +3
Banks analyse sequential financial transaction data to perform many tasks, including fraud prevention, credit risk assessment and offer personalization. To improve the predictive a…
Corruption Robust Offline Reinforcement Learning with Human Feedback
Debmalya Mandal, Andi Nika, Parameswaran Kamalaruban +2
We study data corruption robustness for reinforcement learning with human feedback (RLHF) in an offline setting. Given an offline dataset of pairs of trajectories along with feedba…
DMAP: A Distribution Map for Text
Tom Kempton, Julia Rozanova, Parameswaran Kamalaruban +5
Large Language Models (LLMs) are a powerful tool for statistical text analysis, with derived sequences of next-token probability distributions offering a wealth of information. Ext…
Corruption-robust Offline Multi-agent Reinforcement Learning From Human Feedback
Andi Nika, Debmalya Mandal, Parameswaran Kamalaruban +2
We consider robustness against data corruption in offline multi-agent reinforcement learning from human feedback (MARLHF) under a strong-contamination model: given a dataset of…
Emergent Bias and Fairness in Multi-Agent Decision Systems
Maeve Madigan, Parameswaran Kamalaruban, Glenn Moynihan +3
Multi-agent systems have demonstrated the ability to improve performance on a variety of predictive tasks by leveraging collaborative decision making. However, the lack of effectiv…
Inference-Time Personalized Alignment with a Few User Preference Queries
Victor-Alexandru PÄdurean, Parameswaran Kamalaruban, Nachiket Kotalwar +2
We study the problem of aligning a generative model's response with a user's preferences. Recent works have proposed several different formulations for personalized alignment; howe…