8 papers
VFUSE: Virulent Feature Understanding with Sparse autoEncoders
Michael Yu, Matthew L. Olson
Generative models have shown remarkable progress in a variety of domains such as protein design, but such power enables the opaque generation of hazardous proteins. In this work, w…
LieCraft: A Multi-Agent Framework for Evaluating Deceptive Capabilities in Language Models
Matthew Lyle Olson, Neale Ratzlaff, Musashi Hinck +5
Large Language Models (LLMs) exhibit impressive general-purpose capabilities but also introduce serious safety risks, particularly the potential for deception as models acquire inc…
Probing the Representational Power of Sparse Autoencoders in Vision Models
Matthew Lyle Olson, Musashi Hinck, Neale Ratzlaff +4
Sparse Autoencoders (SAEs) have emerged as a popular tool for interpreting the hidden states of large language models (LLMs). By learning to reconstruct activations from a sparse b…
Debias your Large Multi-Modal Model at Test-Time via Non-Contrastive Visual Attribute Steering
Neale Ratzlaff, Matthew Lyle Olson, Musashi Hinck +4
Large Multi-Modal Models (LMMs) have demonstrated impressive capabilities as general-purpose chatbots able to engage in conversations about visual inputs. However, their responses…
Analyzing Hierarchical Structure in Vision Models with Sparse Autoencoders
Matthew Lyle Olson, Musashi Hinck, Neale Ratzlaff +4
The ImageNet hierarchy provides a structured taxonomy of object categories, offering a valuable lens through which to analyze the representations learned by deep vision models. In…
Probing Semantic Routing in Large Mixture-of-Expert Models
Matthew Lyle Olson, Neale Ratzlaff, Musashi Hinck +4
In the past year, large (>100B parameter) mixture-of-expert (MoE) models have become increasingly common in the open domain. While their advantages are often framed in terms of eff…