5 citations · 6 across the 3 of their papers we have counts for
Showing cs.CRShow all
2 papers · 1 filter
cs.CR2025
X-Guard: Multilingual Guard Agent for Content Moderation
Bibek Upadhayay, Vahid Behzadan, Ph. D
Large Language Models (LLMs) have rapidly become integral to numerous applications in critical domains where reliability is paramount. Despite significant advances in safety framew…
cs.CR2024★ 1 cited
Covert Malicious Finetuning: Challenges in Safeguarding LLM Adaptation
Danny Halawi, Alexander Wei, Eric Wallace +3
Black-box finetuning is an emerging interface for adapting state-of-the-art language models to user needs. However, such access may also let malicious actors undermine model safety…