Showing cs.CRShow all
2 papers · 1 filter
cs.CR2025
Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval
Taiye Chen, Zeming Wei, Ang Li +1
Large Language Models (LLMs) are known to be vulnerable to jailbreaking attacks, wherein adversaries exploit carefully engineered prompts to induce harmful or unethical responses.…
cs.CR2025
X-Guard: Multilingual Guard Agent for Content Moderation
Bibek Upadhayay, Vahid Behzadan, Ph. D
Large Language Models (LLMs) have rapidly become integral to numerous applications in critical domains where reliability is paramount. Despite significant advances in safety framew…