11 citations · 14 across the 5 of their papers we have counts for
Showing cs.CLShow all
2 papers · 1 filter
cs.CL2025
Safeguarding Large Language Models in Real-time with Tunable Safety-Performance Trade-offs
Joao Fonseca, Andrew Bell, Julia Stoyanovich
Large Language Models (LLMs) have been shown to be susceptible to jailbreak attacks, or adversarial attacks used to illicit high risk behavior from a model. Jailbreaks have been ex…
cs.CL2024
Output Scouting: Auditing Large Language Models for Catastrophic Responses
Andrew Bell, Joao Fonseca
Recent high profile incidents in which the use of Large Language Models (LLMs) resulted in significant harm to individuals have brought about a growing interest in AI safety. One r…