Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
Adversarial Déjà Vu: Jailbreak Dictionary Learning for Stronger Generalization to Unseen Attacks
Mahavir Dabas, Tran Huynh, Nikhil Reddy Billa +8
Large language models remain vulnerable to jailbreak attacks that bypass safety guardrails to elicit harmful outputs. Defending against novel jailbreaks represents a critical chall…
cs.LG2025
K-Edit: Language Model Editing with Contextual Knowledge Awareness
Elan Markowitz, Anil Ramakrishna, Ninareh Mehrabi +4
As the world changes, we need to be able to update our models and correct false information without costly retraining. Knowledge-based model editing enables precise modifications t…