6 papers
Beyond I'm Sorry, I Can't: Dissecting Large Language Model Refusal
Nirmalendu Prakash, Yeo Wei Jie, Amir Abdullah +3
Refusal on harmful prompts is a key safety behaviour in instruction-tuned large language models (LLMs), yet the internal causes of this behaviour remain poorly understood. We study…
Geometry-Aware CLIP Retrieval via Local Cross-Modal Alignment and Steering
Nirmalendu Prakash, Narmeen Fatimah Oozeer, Xin Su +8
CLIP retrieval is typically framed as a pointwise similarity problem in a shared embedding space. While CLIP achieves strong global cross-modal alignment, many retrieval failures a…
DreamReader: An Interpretability Toolkit for Text-to-Image Models
Nirmalendu Prakash, Narmeen Oozeer, Michael Lan +6
Despite the rapid adoption of text-to-image (T2I) diffusion models, causal and representation-level analysis remains fragmented and largely limited to isolated probing techniques.…
Activation Space Interventions Can Be Transferred Between Large Language Models
Narmeen Oozeer, Dhruv Nathawani, Nirmalendu Prakash +3
The study of representation universality in AI models reveals growing convergence across domains, modalities, and architectures. However, the practical applications of representati…
Distribution-Aware Feature Selection for SAEs
Narmeen Oozeer, Nirmalendu Prakash, Michael Lan +2
Sparse autoencoders (SAEs) decompose neural activations into interpretable features. A widely adopted variant, the TopK SAE, reconstructs each token from its K most active latents.…
Understanding Refusal in Language Models with Sparse Autoencoders
Wei Jie Yeo, Nirmalendu Prakash, Clement Neo +3
Refusal is a key safety behavior in aligned language models, yet the internal mechanisms driving refusals remain opaque. In this work, we conduct a mechanistic study of refusal in…