collaborators

18 papers

cs.LG2026

Model Forensics: Investigating Whether Concerning Behavior Reflects Misalignment

Aditya Singh, Gerson Kroiz, Senthooran Rajamanoharan +1

A central goal of safety research is determining whether a model is misaligned. Prior work has largely focused on detecting concerning behavior. But behavior alone does not establi…

cs.AI2026

Chain-of-Thought Reasoning In The Wild Is Not Always Faithful

Iván Arcuschin, Jett Janiak, Robert Krzyzanowski +3

Recent studies indicate that when faced with explicit biases in prompts, models often omit mentioning these biases in their Chain-of-Thought (CoT) output, revealing that verbalized…

cs.AI2026

Subliminal Learning Is Steering Vector Distillation

Camila Blank, Agam Bhatia, Senthooran Rajamanoharan +2

Subliminal learning refers to a student language model acquiring a teacher's traits (e.g. a system-prompted preference for owls) when fine-tuned on the teacher's outputs, despite t…

cs.AI2026

How Well Do Models Follow Their Constitutions?

Arya Jakkli, Senthooran Rajamanoharan, Neel Nanda

Frontier AI developers now train models against long written behavioral specifications, such as Anthropic's constitution (Anthropic, 2025a) and OpenAI's Model Spec (OpenAI, 2025a),…

cs.LG2026

Thought Branches: Interpreting LLM Reasoning Requires Resampling

Uzay Macar, Paul C. Bogdan, Senthooran Rajamanoharan +1

Most work interpreting reasoning models studies only a single chain-of-thought (CoT), yet these models define distributions over many possible CoTs. We argue that studying a single…

cs.AI2026

Emergent Misalignment is Easy, Narrow Misalignment is Hard

Anna Soligo, Edward Turner, Senthooran Rajamanoharan +1

Finetuning large language models on narrowly harmful datasets can cause them to become emergently misaligned, giving stereotypically `evil' responses across diverse unrelated setti…