1 paper
Ali Dasdan, Manan Shah, W. Russell Neuman +3
Behavioral audits of Large Language Models on moral prompts measure what the model says, not the internal computation producing it. We use Transluce, an AI-driven mechanistic-inter…