11 papers
When Activation Oracles Learn Not to Read: Concept-Specific Blind Spots in Fine-Tuned Oracles
Tobias Bersia, Tatiana Gaintseva
Activation Oracles (AOs) are language models trained to answer natural-language questions about another model's internal activations. They offer a flexible interface for reading hi…
EquiSteer: Cross-Attention Steering Towards a Fairer Text-Guided Image Generation
Tatiana Gaintseva, Akshit Achara, Gregory Slabaugh +2
Text-to-image diffusion models power everyday creative tasks, but they still reproduce the demographic biases in their training data. On common prompts such as ``a photo of a nurse…
A Geometric Account of Activation Steering through Angle-Norm Decomposition
Georgii Aparin, Tatiana Gaintseva
Linear activation steering has gained popularity as a simple and empirically effective way to control language model behavior. More recently, spherical steering paradigms have been…
MidSteer: Optimal Affine Framework for Steering Generative Models
Tatiana Gaintseva, Andrew Stepanov, Ziquan Liu +4
Steering intermediate representations has emerged as a powerful strategy for controlling generative models, particularly in post-deployment alignment and safety settings. However,…
SV-Detect: AI-generated Text Detection with Steering Vectors
Mikhail Vishnyakov, Tatiana Gaintseva
Detecting AI-generated text is especially difficult under distribution shift, such as transfer across domains, source models, and editing attacks. We propose an AI-generated text d…
Multi-Way Representation Alignment
Akshit Achara, Tatiana Gaintseva, Mateo Mahaut +5
The Platonic Representation Hypothesis suggests that independently trained neural networks converge to increasingly similar latent spaces. However, current strategies for mapping t…