collaborators

5 papers

cs.HC2025

Confirmation bias: A challenge for scalable oversight

Gabriel Recchia, Chatrik Singh Mangat, Jinu Nyachhyon +4

Scalable oversight protocols aim to empower evaluators to accurately verify AI models more capable than themselves. However, human evaluators are subject to biases that can lead to…

cs.CY2025

From Stability to Inconsistency: A Study of Moral Preferences in LLMs

Monika Jotautaite, Mary Phuong, Chatrik Singh Mangat +1

As large language models (LLMs) increasingly integrate into our daily lives, it becomes crucial to understand their implicit biases and moral tendencies. To address this, we introd…

cs.AI2025

FindTheFlaws: Annotated Errors for Detecting Flawed Reasoning and Scalable Oversight Research

Gabriel Recchia, Chatrik Singh Mangat, Issac Li +1

As AI models tackle increasingly complex problems, ensuring reliable human oversight becomes more challenging due to the difficulty of verifying solutions. Approaches to scaling AI…

cs.LG2024

Characterizing stable regions in the residual stream of LLMs

Jett Janiak, Jacek Karwowski, Chatrik Singh Mangat +3

We identify stable regions in the residual stream of Transformers, where the model's output remains insensitive to small activation changes, but exhibits high sensitivity at region…

cs.LG2024

Evaluating Synthetic Activations composed of SAE Latents in GPT-2

Giorgi Giglemiani, Nora Petrova, Chatrik Singh Mangat +2

Sparse Auto-Encoders (SAEs) are commonly employed in mechanistic interpretability to decompose the residual stream into monosemantic SAE latents. Recent work demonstrates that pert…