Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
Exposing the Illusion of Erasure in Knowledge Editing for LLMs
Advik Raj Basani, Anshuman Chhabra
Knowledge Editing (KE) has emerged as a frontier for updating specific facts in LLMs without costly retraining, but its reliability and underlying mechanisms remain poorly understo…
cs.LG2025
GASP: Efficient Black-Box Generation of Adversarial Suffixes for Jailbreaking LLMs
Advik Raj Basani, Xiao Zhang
LLMs have shown impressive capabilities across various natural language processing tasks, yet remain vulnerable to input prompts, known as jailbreak attacks, carefully designed to…