7 citations · 12 across the 2 of their papers we have counts for
2 papers
cs.LG2024★ 5 cited
AgentHarm: A Benchmark for Measuring Harmfulness of LLM Agents
Maksym Andriushchenko, Alexandra Souly, Mateusz Dziemian +11
The robustness of LLMs to jailbreak attacks, where users design prompts to circumvent safety measures and misuse model capabilities, has been studied primarily for LLMs acting as s…
cs.LG2024★ 7 cited
Improving Alignment and Robustness with Circuit Breakers
Andy Zou, Long Phan, Justin Wang +7
AI systems can take harmful actions and are highly vulnerable to adversarial attacks. We present an approach, inspired by recent advances in representation engineering, that interr…