6 papers · 1 filter
Probing the Representational Power of Sparse Autoencoders in Vision Models
Matthew Lyle Olson, Musashi Hinck, Neale Ratzlaff +4
Sparse Autoencoders (SAEs) have emerged as a popular tool for interpreting the hidden states of large language models (LLMs). By learning to reconstruct activations from a sparse b…
Debias your Large Multi-Modal Model at Test-Time via Non-Contrastive Visual Attribute Steering
Neale Ratzlaff, Matthew Lyle Olson, Musashi Hinck +4
Large Multi-Modal Models (LMMs) have demonstrated impressive capabilities as general-purpose chatbots able to engage in conversations about visual inputs. However, their responses…
Analyzing Hierarchical Structure in Vision Models with Sparse Autoencoders
Matthew Lyle Olson, Musashi Hinck, Neale Ratzlaff +4
The ImageNet hierarchy provides a structured taxonomy of object categories, offering a valuable lens through which to analyze the representations learned by deep vision models. In…
ClimDetect: A Benchmark Dataset for Climate Change Detection and Attribution
Sungduk Yu, Brian L. White, Anahita Bhiwandiwalla +6
Detecting and attributing temperature increases driven by climate change is crucial for understanding global warming and informing adaptation strategies. However, distinguishing hu…
Debiasing Large Vision-Language Models by Ablating Protected Attribute Representations
Neale Ratzlaff, Matthew Lyle Olson, Musashi Hinck +3
Large Vision Language Models (LVLMs) such as LLaVA have demonstrated impressive capabilities as general-purpose chatbots that can engage in conversations about a provided input ima…
LVLM-Interpret: An Interpretability Tool for Large Vision-Language Models
Gabriela Ben Melech Stan, Estelle Aflalo, Raanan Yehezkel Rohekar +7
In the rapidly evolving landscape of artificial intelligence, multi-modal large language models are emerging as a significant area of interest. These models, which combine various…