collaborators
Showing cs.LGShow all

5 papers · 1 filter

cs.LG2026

Reading the Finetuning Prior: Verbatim Content Recovery via Contrastive Decoding Diffing

Michał Brzozowski, Zuzanna Dubanowska, Enrico Cassano +1

Narrowly finetuned language models memorize implanted content verbatim, but auditing what a deployed model has been taught, without access to its weights or training data, remains…

cs.LG2026

Aligned Training: A Parameter-Free Method to Improve Feature Quality and Stability of Sparse Autoencoders (SAE)

Michał Brzozowski, Neo Christopher Chung

Sparse autoencoders (SAEs) are one of the main methods to interpret the inner workings of deep neural networks (DNNs), decomposing activations into higher-dimensional features. How…

cs.LG2026

Ablating Archetypes: The Stability of Archetypal SAEs is an Artifact of Initialization and Metric Design

Michał Brzozowski, Neo Christopher Chung

Dictionary learning with sparse autoencoders (SAEs) produces overcomplete bases from neural network activations that are often interpretable and reduces polysemanticity. However, f…

cs.LG2026

GPart: End-to-End Isometric Fine-Tuning via Global Parameter Partitioning

Paolo Mandica, Michał Brzozowski, Zuzanna Dubanowska +1

Low-rank adaptation (LoRA) has become the dominant paradigm for parameter-efficient fine-tuning (PEFT) of large language models (LLMs). However, its bilinear structure introduces a…

cs.LG2025

Representation-based Broad Hallucination Detectors Fail to Generalize Out of Distribution

Zuzanna Dubanowska, Maciej Żelaszczyk, Michał Brzozowski +2

We critically assess the efficacy of the current SOTA in hallucination detection and find that its performance on the RAGTruth dataset is largely driven by a spurious correlation w…