collaborators

5 papers

cs.CL2026

Reasoning Fine-Tuning Induces Persistent Latent Policy States

Abir Harrasse, Michael Lan, Hunar Batra +2

Reasoning-specialized language models show large performance gains over base models, yet the internal changes responsible for improved multi-step reasoning remain poorly understood…

cs.CY2026

Make Mechanistic Interpretability Auditable: A Call to Develop Guidelines via Continuous Collaborative Reviewing

Michael Lan, Narmeen Fatimah Oozeer, Chaithanya Bandi +4

While mechanistic interpretability (MI) has produced important insights into neural network internals, the field has yet to establish a standardized system to audit experiments. As…

cs.LG2026

DreamReader: An Interpretability Toolkit for Text-to-Image Models

Nirmalendu Prakash, Narmeen Oozeer, Michael Lan +6

Despite the rapid adoption of text-to-image (T2I) diffusion models, causal and representation-level analysis remains fragmented and largely limited to isolated probing techniques.…

cs.AI2025

Activation Space Interventions Can Be Transferred Between Large Language Models

Narmeen Oozeer, Dhruv Nathawani, Nirmalendu Prakash +3

The study of representation universality in AI models reveals growing convergence across domains, modalities, and architectures. However, the practical applications of representati…

cs.LG2025

Distribution-Aware Feature Selection for SAEs

Narmeen Oozeer, Nirmalendu Prakash, Michael Lan +2

Sparse autoencoders (SAEs) decompose neural activations into interpretable features. A widely adopted variant, the TopK SAE, reconstructs each token from its K most active latents.…