1 citations · 1 across the 11 of their papers we have counts for
15 papers · 1 filter
Corruption Robust Offline Reinforcement Learning with Human Feedback
Debmalya Mandal, Andi Nika, Parameswaran Kamalaruban +2
We study data corruption robustness for reinforcement learning with human feedback (RLHF) in an offline setting. Given an offline dataset of pairs of trajectories along with feedba…
MaMa: A Game-Theoretic Approach for Designing Safe Agentic Systems
Jonathan Nöther, Adish Singla, Goran Radanovic
LLM-based multi-agent systems have demonstrated impressive capabilities, but they also introduce significant safety risks when individual agents fail or behave adversarially. In th…
Corruption-robust Offline Multi-agent Reinforcement Learning From Human Feedback
Andi Nika, Debmalya Mandal, Parameswaran Kamalaruban +2
We consider robustness against data corruption in offline multi-agent reinforcement learning from human feedback (MARLHF) under a strong-contamination model: given a dataset of…
Learning Half-Spaces from Perturbed Contrastive Examples
Aryan Alavi Razavi Ravari, Farnam Mansouri, Yuxin Chen +3
We study learning under a two-step contrastive example oracle, as introduced by Mansouri et. al. (2025), where each queried (or sampled) labeled example is paired with an additiona…
Inference-Time Personalized Alignment with a Few User Preference Queries
Victor-Alexandru PÄdurean, Parameswaran Kamalaruban, Nachiket Kotalwar +2
We study the problem of aligning a generative model's response with a user's preferences. Recent works have proposed several different formulations for personalized alignment; howe…
Curriculum Design for Trajectory-Constrained Agent: Compressing Chain-of-Thought Tokens in LLMs
Georgios Tzannetos, Parameswaran Kamalaruban, Adish Singla
Training agents to operate under strict constraints during deployment, such as limited resource budgets or stringent safety requirements, presents significant challenges, especiall…