2 papers
cs.CL2025
SHIELD: Classifier-Guided Prompting for Robust and Safer LVLMs
Juan Ren, Mark Dras, Usman Naseem
Large Vision-Language Models (LVLMs) unlock powerful multimodal reasoning but also expand the attack surface, particularly through adversarial inputs that conceal harmful goals in…
cs.CL2025
Steering Over-refusals Towards Safety in Retrieval Augmented Generation
Utsav Maskey, Mark Dras, Usman Naseem
Safety alignment in large language models (LLMs) induces over-refusals -- where LLMs decline benign requests due to aggressive safety filters. We analyze this phenomenon in retriev…