4 citations · 5 across the 3 of their papers we have counts for
Showing cs.LGShow all
2 papers · 1 filter
cs.LG2025★ 1 cited
DynaGuard: A Dynamic Guardian Model With User-Defined Policies
Monte Hoover, Vatsal Baherwani, Neel Jain +7
Guardian models play a crucial role in ensuring the safety and ethical behavior of user-facing AI applications by enforcing guardrails and detecting harmful content. While standard…
cs.LG2024★ 4 cited
Coercing LLMs to do and reveal (almost) anything
Jonas Geiping, Alex Stein, Manli Shu +3
It has recently been shown that adversarial attacks on large language models (LLMs) can "jailbreak" the model into making harmful statements. In this work, we argue that the spectr…