3 papers
cs.LG2026
Dynamic Weight Grafting: Localizing Finetuned Factual Knowledge in Transformers
Todd Nief, David Reber, Sean Richardson +1
When an LLM learns a new fact during finetuning (e.g., new movie releases, newly elected pope, etc.), where does this information go? Are entities enriched with relation informatio…
cs.CL2025
Incorporating Hierarchical Semantics in Sparse Autoencoder Architectures
Mark Muchane, Sean Richardson, Kiho Park +1
Sparse dictionary learning (and, in particular, sparse autoencoders) attempts to learn a set of human-understandable concepts that can explain variation on an abstract space. A bas…
cs.CL2025
RATE: Causal Explainability of Reward Models with Imperfect Counterfactuals
David Reber, Sean Richardson, Todd Nief +2
Reward models are widely used as proxies for human preferences when aligning or evaluating LLMs. However, reward models are black boxes, and it is often unclear what, exactly, they…