5 papers · 1 filter
Reading the Finetuning Prior: Verbatim Content Recovery via Contrastive Decoding Diffing
MichaÅ Brzozowski, Zuzanna Dubanowska, Enrico Cassano +1
Narrowly finetuned language models memorize implanted content verbatim, but auditing what a deployed model has been taught, without access to its weights or training data, remains…
Aligned Training: A Parameter-Free Method to Improve Feature Quality and Stability of Sparse Autoencoders (SAE)
MichaÅ Brzozowski, Neo Christopher Chung
Sparse autoencoders (SAEs) are one of the main methods to interpret the inner workings of deep neural networks (DNNs), decomposing activations into higher-dimensional features. How…
Ablating Archetypes: The Stability of Archetypal SAEs is an Artifact of Initialization and Metric Design
MichaÅ Brzozowski, Neo Christopher Chung
Dictionary learning with sparse autoencoders (SAEs) produces overcomplete bases from neural network activations that are often interpretable and reduces polysemanticity. However, f…
GPart: End-to-End Isometric Fine-Tuning via Global Parameter Partitioning
Paolo Mandica, MichaÅ Brzozowski, Zuzanna Dubanowska +1
Low-rank adaptation (LoRA) has become the dominant paradigm for parameter-efficient fine-tuning (PEFT) of large language models (LLMs). However, its bilinear structure introduces a…
Representation-based Broad Hallucination Detectors Fail to Generalize Out of Distribution
Zuzanna Dubanowska, Maciej Żelaszczyk, MichaŠBrzozowski +2
We critically assess the efficacy of the current SOTA in hallucination detection and find that its performance on the RAGTruth dataset is largely driven by a spurious correlation w…