1 citations · 1 across the 5 of their papers we have counts for
Showing 2024Show all
2 papers · 1 filter
cs.CL2024★ 1 cited
AmpleGCG-Plus: A Strong Generative Model of Adversarial Suffixes to Jailbreak LLMs with Higher Success Rates in Fewer Attempts
Vishal Kumar, Zeyi Liao, Jaylen Jones +1
Although large language models (LLMs) are typically aligned, they remain vulnerable to jailbreaking through either carefully crafted prompts in natural language or, interestingly,…
cs.CL2024
A Multi-Aspect Framework for Counter Narrative Evaluation using Large Language Models
Jaylen Jones, Lingbo Mo, Eric Fosler-Lussier +1
Counter narratives - informed responses to hate speech contexts designed to refute hateful claims and de-escalate encounters - have emerged as an effective hate speech intervention…