13 papers
Meta-Learning Preferences for Multilingual LLM Alignment
Jiaying Lin, Seongho Son, Nam Phuong Tran +3
The paper introduces a meta-learning method that uses preference data from high-resource languages to quickly adapt large language models to low-resource languages with very few hu…
Corruption Robust Offline Reinforcement Learning with Human Feedback
Debmalya Mandal, Andi Nika, Parameswaran Kamalaruban +2
We study data corruption robustness for reinforcement learning with human feedback (RLHF) in an offline setting. Given an offline dataset of pairs of trajectories along with feedba…
Distributionally Robust Reinforcement Learning with Human Feedback
Debmalya Mandal, Paulius Sasnauskas, Goran Radanovic
Reinforcement learning from human feedback (RLHF) has evolved to be one of the main methods for fine-tuning large language models (LLMs). However, existing RLHF methods are non-rob…
Sparse Offline Reinforcement Learning with Corruption Robustness
Nam Phuong Tran, Andi Nika, Goran Radanovic +2
We investigate robustness to strong data corruption in offline sparse reinforcement learning (RL). In our setting, an adversary may arbitrarily perturb a fraction of the collected…
GRAPHLCP: Structure-Aware Localized Conformal Prediction on Graphs
Peyman Baghershahi, Fangxin Wang, Debmalya Mandal +1
Conformal prediction (CP) provides a distribution-free approach to uncertainty quantification with finite-sample guarantees. However, applying CP to graph neural networks (GNNs) re…
Corruption-robust Offline Multi-agent Reinforcement Learning From Human Feedback
Andi Nika, Debmalya Mandal, Parameswaran Kamalaruban +2
We consider robustness against data corruption in offline multi-agent reinforcement learning from human feedback (MARLHF) under a strong-contamination model: given a dataset of…