activity
20242026
most citedCorruption-Robust Offline Two-Player Zero-Sum Markov Games

1 citations · 1 across the 3 of their papers we have counts for

collaborators

6 papers

cs.LG2026

Corruption-robust Offline Multi-agent Reinforcement Learning From Human Feedback

Andi Nika, Debmalya Mandal, Parameswaran Kamalaruban +2

We consider robustness against data corruption in offline multi-agent reinforcement learning from human feedback (MARLHF) under a strong-contamination model: given a dataset of…

stat.ML2025

Sparse Offline Reinforcement Learning with Corruption Robustness

Nam Phuong Tran, Andi Nika, Goran Radanovic +2

We investigate robustness to strong data corruption in offline sparse reinforcement learning (RL). In our setting, an adversary may arbitrarily perturb a fraction of the collected…

cs.LG2025

Policy Teaching via Data Poisoning in Learning from Human Preferences

Andi Nika, Jonathan Nöther, Debmalya Mandal +3

We study data poisoning attacks in learning from human preferences. More specifically, we consider the problem of teaching/enforcing a target policy by synthesizing pre…

cs.GT20241 cited

Corruption-Robust Offline Two-Player Zero-Sum Markov Games

Andi Nika, Debmalya Mandal, Adish Singla +1

We study data corruption robustness in offline two-player zero-sum Markov games. Given a dataset of realized trajectories of two players, an adversary is allowed to modify an -f…

cs.LG2024

Reward Model Learning vs. Direct Policy Optimization: A Comparative Analysis of Learning from Human Preferences

Andi Nika, Debmalya Mandal, Parameswaran Kamalaruban +3

In this paper, we take a step towards a deeper understanding of learning from human preferences by systematically comparing the paradigm of reinforcement learning from human feedba…

cs.LG2024

Corruption Robust Offline Reinforcement Learning with Human Feedback

Debmalya Mandal, Andi Nika, Parameswaran Kamalaruban +2

We study data corruption robustness for reinforcement learning with human feedback (RLHF) in an offline setting. Given an offline dataset of pairs of trajectories along with feedba…