collaborators

7 papers

cs.CL2026

Gotta Catch them all: the modes of Sycophancy

Shreyans Jain, Alexandra Yost, Amirali Abdullah

Large language models often align with users' beliefs at the expense of factual accuracy, a behavior known as sycophancy. Prior mechanistic studies largely treat sycophancy as a si…

cs.AI2026

The Capability Frontier: Benchmarks Miss 82% of Model Performance

Bradley Fowler, Ryan Smith, Daniel Thi Graviet +8

Existing benchmarks typically report accuracy for a single model on a single run. This systematically understates real-world LLM capabilities, particularly under heterogeneous data…

cs.LG2026

Riemannian-Manifold Steering: Geometry-Aware Generative Autoencoders for Label-Free Steering

Narmeen Oozeer, Shivam Raval, Philip Quirke +4

Steering a language model - intervening on its internal activations to change downstream behaviour - has recently expanded beyond linear interpolation to nonlinear methods such as…

cs.LG2026

STRIDE: Training Data Attribution via Sparse Recovery from Subset Perturbations

Rishit Dagli, Abir Harrasse, Luke Zhang +4

Training Data Attribution (TDA) seeks to trace a model's predictions back to its training data. The gold standard for TDA relies on causal interventions, observing how a model chan…

cs.AI2026

Measure what Matters: Psychometric Evaluation of AI with Situational Judgment Tests

Alexandra Yost, Shreyans Jain, Shivam Raval +6

Persona conditioning is widely used to steer large language model (LLM) behavior, but it is unclear whether it induces stable behavioral structure or superficial variation. We prop…

cs.CY2026

Make Mechanistic Interpretability Auditable: A Call to Develop Guidelines via Continuous Collaborative Reviewing

Michael Lan, Narmeen Fatimah Oozeer, Chaithanya Bandi +4

While mechanistic interpretability (MI) has produced important insights into neural network internals, the field has yet to establish a standardized system to audit experiments. As…