activity
20242026
collaborators

5 papers

cs.AI2026

Every Eval Ever: A Unifying Schema and Community Repository for AI Evaluation Results

Jan Batzner, Sree Harsha Nelaturu, Damian Stachura +45

AI evaluations are widely used for testing and understanding progress. However, the diverse evaluators bring with them inconsistencies that challenge analysis and comparison. First…

cs.CY2026

Make Mechanistic Interpretability Auditable: A Call to Develop Guidelines via Continuous Collaborative Reviewing

Michael Lan, Narmeen Fatimah Oozeer, Chaithanya Bandi +4

While mechanistic interpretability (MI) has produced important insights into neural network internals, the field has yet to establish a standardized system to audit experiments. As…

cs.LG2025

Quantifying Feature Space Universality Across Large Language Models via Sparse Autoencoders

Michael Lan, Philip Torr, Austin Meek +3

The Universality Hypothesis in large language models (LLMs) claims that different models converge towards similar concept representations in their latent spaces. Providing evidence…

cs.AI2025

Auditing language models for hidden objectives

Samuel Marks, Johannes Treutlein, Trenton Bricken +32

We study the feasibility of conducting alignment audits: investigations into whether models have undesired objectives. As a testbed, we train a language model with a hidden objecti…

cs.AI2024

Inducing Human-like Biases in Moral Reasoning Language Models

Artem Karpov, Seong Hah Cho, Austin Meek +3

In this work, we study the alignment (BrainScore) of large language models (LLMs) fine-tuned for moral reasoning on behavioral data and/or brain data of humans performing the same…