4 papers
MaMa: A Game-Theoretic Approach for Designing Safe Agentic Systems
Jonathan Nöther, Adish Singla, Goran Radanovic
LLM-based multi-agent systems have demonstrated impressive capabilities, but they also introduce significant safety risks when individual agents fail or behave adversarially. In th…
Benchmarking the Robustness of Agentic Systems to Adversarially-Induced Harms
Jonathan Nöther, Jonathan Nöther, Adish Singla +1
Ensuring the safe use of agentic systems requires a thorough understanding of the range of malicious behaviors these systems may exhibit. In this paper, we evaluate the robustness…
Policy Teaching via Data Poisoning in Learning from Human Preferences
Andi Nika, Jonathan Nöther, Debmalya Mandal +3
We study data poisoning attacks in learning from human preferences. More specifically, we consider the problem of teaching/enforcing a target policy by synthesizing pr…
Text-Diffusion Red-Teaming of Large Language Models: Unveiling Harmful Behaviors with Proximity Constraints
Jonathan Nöther, Adish Singla, Goran RadanoviÄ
Recent work has proposed automated red-teaming methods for testing the vulnerabilities of a given target large language model (LLM). These methods use red-teaming LLMs to uncover i…