From the 1 of 6 linked papers with an AI index.
6 papers
Value Leakage: An LLM's Answers Are Silently Shaped by Its Own Values
Jan Betley, Johannes Treutlein, Jan DubiÅski +7
The paper identifies and measures covert value leakage, where large language models let their own values subtly bias answers without informing users, and introduces evaluation suit…
Negation Neglect: When models fail to learn negations in training
Harry Mayne, Lev McKinney, Jan DubiÅski +3
We introduce Negation Neglect, where finetuning LLMs on documents that flag a claim as false makes them believe the claim is true. For example, models are finetuned on documents th…
Jailbreaking Vision-Language Models Through the Visual Modality
Aharon Azulay, Jan DubiÅski, Zhuoyun Li +2
The visual modality of vision-language models (VLMs) is an underexplored attack surface for bypassing safety alignment. We introduce four jailbreak attacks exploiting the vision co…
Conditional misalignment: common interventions can hide emergent misalignment behind contextual triggers
Jan DubiÅski, Jan Betley, Anna Sztyber-Betley +2
Finetuning a language model can lead to emergent misalignment (EM) [Betley et al., 2025b]. Models trained on a narrow distribution of misaligned behavior generalize to more egregio…
Conditioned Activation Transport for T2I Safety Steering
Maciej ChrabÄ szcz, Aleksander Szymczyk, Jan DubiÅski +3
Despite their impressive capabilities, current Text-to-Image (T2I) models remain prone to generating unsafe and toxic content. While activation steering offers a promising inferenc…
Learning Graph Representation of Agent Diffusers
Youcef Djenouri, Nassim Belmecheri, Tomasz Michalak +3
Diffusion-based generative models have significantly advanced text-to-image synthesis, demonstrating impressive text comprehension and zero-shot generalization. These models refine…