3 papers
cs.CR2025
Mitigating Jailbreaks with Intent-Aware LLMs
Wei Jie Yeo, Ranjan Satapathy, Erik Cambria
Despite extensive safety-tuning, large language models (LLMs) remain vulnerable to jailbreak attacks via adversarially crafted instructions, reflecting a persistent trade-off betwe…
cs.CL2025
Understanding Refusal in Language Models with Sparse Autoencoders
Wei Jie Yeo, Nirmalendu Prakash, Clement Neo +3
Refusal is a key safety behavior in aligned language models, yet the internal mechanisms driving refusals remain opaque. In this work, we conduct a mechanistic study of refusal in…
cs.CV2025
Debiasing CLIP: Interpreting and Correcting Bias in Attention Heads
Wei Jie Yeo, Rui Mao, Moloud Abdar +2
Multimodal models like CLIP have gained significant attention due to their remarkable zero-shot performance across various tasks. However, studies have revealed that CLIP can inadv…