Showing cs.CLShow all
2 papers · 1 filter
cs.CL2026
ConceptGuard: Benchmarking Context-Sensitive Unlearning in Large Language Models
Sahil Kale, Ian Harris
Large Language Models (LLMs) increasingly require selective removal of harmful or sensitive knowledge, called unlearning, yet existing methods and benchmarks fail to evaluate this…
cs.CL2023
Robust Safety Classifier for Large Language Models: Adversarial Prompt Shield
Jinhwa Kim, Ali Derakhshan, Ian G. Harris
Large Language Models' safety remains a critical concern due to their vulnerability to adversarial attacks, which can prompt these systems to produce harmful responses. In the hear…