4 papers
ConceptGuard: Benchmarking Context-Sensitive Unlearning in Large Language Models
Sahil Kale, Ian Harris
Large Language Models (LLMs) increasingly require selective removal of harmful or sensitive knowledge, called unlearning, yet existing methods and benchmarks fail to evaluate this…
Memisis: Orchestrating and Evaluating Synthetic Data for Tabular Health Datasets
Nitish Nagesh, Pengbao Zhou, Atchuth Naveen Chilaparasetti +8
Synthetic data is widely used in healthcare to create datasets that preserve statistical properties of real data without exposing sensitive patient information. Generating and eval…
Context Misleads LLMs: The Role of Context Filtering in Maintaining Safe Alignment of LLMs
Jinhwa Kim, Ian G. Harris
While Large Language Models (LLMs) have shown significant advancements in performance, various jailbreak attacks have posed growing safety and ethical risks. Malicious users often…
Robust Safety Classifier for Large Language Models: Adversarial Prompt Shield
Jinhwa Kim, Ali Derakhshan, Ian G. Harris
Large Language Models' safety remains a critical concern due to their vulnerability to adversarial attacks, which can prompt these systems to produce harmful responses. In the hear…