collaborators

8 papers

cs.CL2026

Beyond I'm Sorry, I Can't: Dissecting Large Language Model Refusal

Nirmalendu Prakash, Yeo Wei Jie, Amir Abdullah +3

Refusal on harmful prompts is a key safety behaviour in instruction-tuned large language models (LLMs), yet the internal causes of this behaviour remain poorly understood. We study…

cs.CV2026

Geometry-Aware CLIP Retrieval via Local Cross-Modal Alignment and Steering

Nirmalendu Prakash, Narmeen Fatimah Oozeer, Xin Su +8

CLIP retrieval is typically framed as a pointwise similarity problem in a shared embedding space. While CLIP achieves strong global cross-modal alignment, many retrieval failures a…

cs.AI2026

Curveball Steering: The Right Direction To Steer Isn't Always Linear

Shivam Raval, Hae Jin Song, Linlin Wu +4

Activation steering is a widely used approach for controlling large language model (LLM) behavior by intervening on internal representations. Existing methods largely rely on the L…

cs.LG2026

Understanding and Mitigating Dataset Corruption in LLM Steering

Cullen Anderson, Narmeen Oozeer, Foad Namjoo +3

Contrastive steering has been shown as a simple and effective method to adjust the generative behavior of LLMs at inference time. It uses examples of prompt responses with and with…

cs.LG2026

DreamReader: An Interpretability Toolkit for Text-to-Image Models

Nirmalendu Prakash, Narmeen Oozeer, Michael Lan +6

Despite the rapid adoption of text-to-image (T2I) diffusion models, causal and representation-level analysis remains fragmented and largely limited to isolated probing techniques.…

cs.LG2025

TinySQL: A Progressive Text-to-SQL Dataset for Mechanistic Interpretability Research

Abir Harrasse, Philip Quirke, Clement Neo +3

Mechanistic interpretability research faces a gap between analyzing simple circuits in toy tasks and discovering features in large models. To bridge this gap, we propose text-to-SQ…