2 papers
cs.CL2026
What Makes Two Language Models Think Alike?
Louis Jalouzot, Christophe Pallier, Emmanuel Chemla +1
Do architectural and training differences influence the way models represent and process language? Traditional similarity metrics tell us whether two models share a similar represe…
cs.AI2025
Auditing language models for hidden objectives
Samuel Marks, Johannes Treutlein, Trenton Bricken +32
We study the feasibility of conducting alignment audits: investigations into whether models have undesired objectives. As a testbed, we train a language model with a hidden objecti…