4 papers
Status Association Does Not Reliably Predict Decision Leakage
Abdullah X
Bias evaluations often move too quickly from evidence that a model encodes a social association to claims that the same association will alter consequential decisions. We test whet…
Multi-Agent AI Safety as an Institutional Design Problem
Abdullah X
AI agents increasingly work inside systems that govern how they delegate tasks, move information, execute actions, and use shared resources. Recent work already shows that deployme…
Markovian Circuit Tracing for Transformer State Dynamic
Abdullah X
Many sequence computations are easier to study as movement through internal states than as isolated local circuits. We introduce Markovian Circuit Tracing (MCT), a diagnostic pipel…
Unlearning at Scale: State-Exact Trace-Preserving Deletion in Billion-Parameter Language Models
Abdullah X
Can a prospectively instrumented training continuation reproduce a deletion counterfactual exactly after selected examples leave its replay dataset? We study a trace-preserving cou…