4 papers
The Ghost Couple: Correlated LLM Name Priors and Their Haunting of the Web and Academic Publishing
Michał Brzozowski, Neo Christopher Chung
These names do not exist. Elena Vasquez and Marcus Chen have appeared as volcano experts, astronauts, thriller protagonists, podcast hosts, and academic co-authors across hundreds…
Reading the Finetuning Prior: Verbatim Content Recovery via Contrastive Decoding Diffing
Michał Brzozowski, Zuzanna Dubanowska, Enrico Cassano +1
Narrowly finetuned language models memorize implanted content verbatim, but auditing what a deployed model has been taught, without access to its weights or training data, remains…
Aligned Training: A Parameter-Free Method to Improve Feature Quality and Stability of Sparse Autoencoders (SAE)
Michał Brzozowski, Neo Christopher Chung
Sparse autoencoders (SAEs) are one of the main methods to interpret the inner workings of deep neural networks (DNNs), decomposing activations into higher-dimensional features. How…
Representation-based Broad Hallucination Detectors Fail to Generalize Out of Distribution
Zuzanna Dubanowska, Maciej Żelaszczyk, Michał Brzozowski +2
We critically assess the efficacy of the current SOTA in hallucination detection and find that its performance on the RAGTruth dataset is largely driven by a spurious correlation w…