6 papers
Censored LLMs as a Natural Testbed for Secret Knowledge Elicitation
Helena Casademunt, Bartosz CywiÅski, Khoi Tran +3
Large language models sometimes produce false or misleading responses. Two approaches to this problem are honesty elicitation -- modifying prompts or weights so that the model answ…
Precise Parameter Localization for Textual Generation in Diffusion Models
Åukasz Staniszewski, Bartosz CywiÅski, Franziska Boenisch +2
Novel diffusion models can synthesize photo-realistic images with integrated high-quality text. Surprisingly, we demonstrate through attention activation patching that only less th…
Simple LLM Baselines are Competitive for Model Diffing
Elias Kempf, Simon Schrodi, Bartosz CywiÅski +3
Standard LLM evaluations only test capabilities or dispositions that evaluators designed them for, missing unexpected differences such as behavioral shifts between model revisions…
Eliciting Secret Knowledge from Language Models
Bartosz CywiÅski, Emil Ryd, Rowan Wang +4
We study secret elicitation: discovering knowledge that an AI possesses but does not explicitly verbalize. As a testbed, we train three families of large language models (LLMs) to…
SAeUron: Interpretable Concept Unlearning in Diffusion Models with Sparse Autoencoders
Bartosz CywiÅski, Kamil Deja
Diffusion models, while powerful, can inadvertently generate harmful or undesirable content, raising significant ethical and safety concerns. Recent machine unlearning approaches o…
Towards eliciting latent knowledge from LLMs with mechanistic interpretability
Bartosz CywiÅski, Emil Ryd, Senthooran Rajamanoharan +1
As language models become more powerful and sophisticated, it is crucial that they remain trustworthy and reliable. There is concerning preliminary evidence that models may attempt…