5 papers
Transformers learn factored representations
Adam Shai, Loren Amdahl-Culleton, Casper L. Christensen +6
Transformers pretrained via next token prediction learn to factor their world into parts, representing these factors in orthogonal subspaces of the residual stream. We formalize tw…
Decomposition of Small Transformer Models
Casper L. Christensen, Logan Riggs
Recent work in mechanistic interpretability has shown that decomposing models in parameter space may yield clean handles for analysis and intervention. Previous methods have demons…
Code Like Humans: A Multi-Agent Solution for Medical Coding
Andreas Motzfeldt, Joakim Edin, Casper L. Christensen +3
In medical coding, experts map unstructured clinical notes to alphanumeric codes for diagnoses and procedures. We introduce Code Like Humans: a new agentic framework for medical co…
Bi-Axial Transformers: Addressing the Increasing Complexity of EHR Classification
Rachael DeVries, Casper Christensen, Marie Lisandra Zepeda Mendoza +1
Electronic Health Records (EHRs), the digital representation of a patient's medical history, are a valuable resource for epidemiological and clinical research. They are also becomi…
Normalized AOPC: Fixing Misleading Faithfulness Metrics for Feature Attribution Explainability
Joakim Edin, Andreas Geert Motzfeldt, Casper L. Christensen +3
Deep neural network predictions are notoriously difficult to interpret. Feature attribution methods aim to explain these predictions by identifying the contribution of each input f…