2 citations · 2 across the 2 of their papers we have counts for
Showing cs.LGShow all
2 papers · 1 filter
cs.LG2025
Text-Diffusion Red-Teaming of Large Language Models: Unveiling Harmful Behaviors with Proximity Constraints
Jonathan Nöther, Adish Singla, Goran Radanović
Recent work has proposed automated red-teaming methods for testing the vulnerabilities of a given target large language model (LLM). These methods use red-teaming LLMs to uncover i…
cs.LG2023★ 2 cited
Implicit Poisoning Attacks in Two-Agent Reinforcement Learning: Adversarial Policies for Training-Time Attacks
Mohammad Mohammadi, Jonathan Nöther, Debmalya Mandal +2
In targeted poisoning attacks, an attacker manipulates an agent-environment interaction to force the agent into adopting a policy of interest, called target policy. Prior work has…