2 papers
cs.CL2026
Most Current Model Organisms Are Leaky: Perplexity Differencing Often Reveals Finetuning Objectives
Mohammed Abu Baker, Luca Baroni, Dan Wilhelm
Finetuning can significantly modify the behavior of large language models, including introducing harmful or unsafe behaviors. To study these risks, researchers develop model organi…
cs.LG2025
Tokenized SAEs: Disentangling SAE Reconstructions
Thomas Dooms, Daniel Wilhelm
Sparse auto-encoders (SAEs) have become a prevalent tool for interpreting language models' inner workings. However, it is unknown how tightly SAE features correspond to computation…