2 papers
cs.CR2025
Context Misleads LLMs: The Role of Context Filtering in Maintaining Safe Alignment of LLMs
Jinhwa Kim, Ian G. Harris
While Large Language Models (LLMs) have shown significant advancements in performance, various jailbreak attacks have posed growing safety and ethical risks. Malicious users often…
cs.CL2023
Robust Safety Classifier for Large Language Models: Adversarial Prompt Shield
Jinhwa Kim, Ali Derakhshan, Ian G. Harris
Large Language Models' safety remains a critical concern due to their vulnerability to adversarial attacks, which can prompt these systems to produce harmful responses. In the hear…