44 citations · 85 across the 14 of their papers we have counts for
19 papers · 1 filter
MaMa: A Game-Theoretic Approach for Designing Safe Agentic Systems
Jonathan Nöther, Adish Singla, Goran Radanovic
LLM-based multi-agent systems have demonstrated impressive capabilities, but they also introduce significant safety risks when individual agents fail or behave adversarially. In th…
Benchmarking the Robustness of Agentic Systems to Adversarially-Induced Harms
Jonathan Nöther, Adish Singla, Goran Radanovic
Ensuring the safe use of agentic systems requires a thorough understanding of the range of malicious behaviors these systems may exhibit. In this paper, we evaluate the robustness…
On Corruption-Robustness in Performative Reinforcement Learning
Vasilis Pollatos, Debmalya Mandal, Goran Radanovic
In performative Reinforcement Learning (RL), an agent faces a policy-dependent environment: the reward and transition functions depend on the agent's policy. Prior work on performa…
Policy Teaching via Data Poisoning in Learning from Human Preferences
Andi Nika, Jonathan Nöther, Debmalya Mandal +3
We study data poisoning attacks in learning from human preferences. More specifically, we consider the problem of teaching/enforcing a target policy by synthesizing pre…
Distributionally Robust Reinforcement Learning with Human Feedback
Debmalya Mandal, Paulius Sasnauskas, Goran Radanovic
Reinforcement learning from human feedback (RLHF) has evolved to be one of the main methods for fine-tuning large language models (LLMs). However, existing RLHF methods are non-rob…
Text-Diffusion Red-Teaming of Large Language Models: Unveiling Harmful Behaviors with Proximity Constraints
Jonathan Nöther, Adish Singla, Goran Radanović
Recent work has proposed automated red-teaming methods for testing the vulnerabilities of a given target large language model (LLM). These methods use red-teaming LLMs to uncover i…