2 papers
cs.LG2026
PLOT: Progressive Localization via Optimal Transport in Neural Causal Abstraction
Jonathn Chang, Arya Datla, Ziv Goldfeld
Causal abstraction offers a principled framework for mechanistic interpretability, aligning a high-level causal model with the low-level computation realized by a neural network th…
cs.AI2026
EigenBench: A Comparative Behavioral Measure of Value Alignment
Jonathn Chang, Leonhard Piff, Suvadip Sana +2
Aligning AI with human values is a pressing unsolved problem. To address the lack of quantitative metrics for value alignment, we propose EigenBench: a black-box method for compara…