activity
20242026
collaborators

11 papers

cs.AI2026

LieCraft: A Multi-Agent Framework for Evaluating Deceptive Capabilities in Language Models

Matthew Lyle Olson, Neale Ratzlaff, Musashi Hinck +5

Large Language Models (LLMs) exhibit impressive general-purpose capabilities but also introduce serious safety risks, particularly the potential for deception as models acquire inc…

cs.CV2025

Probing the Representational Power of Sparse Autoencoders in Vision Models

Matthew Lyle Olson, Musashi Hinck, Neale Ratzlaff +4

Sparse Autoencoders (SAEs) have emerged as a popular tool for interpreting the hidden states of large language models (LLMs). By learning to reconstruct activations from a sparse b…

cs.CV2025

Debias your Large Multi-Modal Model at Test-Time via Non-Contrastive Visual Attribute Steering

Neale Ratzlaff, Matthew Lyle Olson, Musashi Hinck +4

Large Multi-Modal Models (LMMs) have demonstrated impressive capabilities as general-purpose chatbots able to engage in conversations about visual inputs. However, their responses…

cs.AI2025

DPO Learning with LLMs-Judge Signal for Computer Use Agents

Man Luo, David Cobbley, Xin Su +4

Computer use agents (CUA) are systems that automatically interact with graphical user interfaces (GUIs) to complete tasks. CUA have made significant progress with the advent of lar…

cs.CV2025

Analyzing Hierarchical Structure in Vision Models with Sparse Autoencoders

Matthew Lyle Olson, Musashi Hinck, Neale Ratzlaff +4

The ImageNet hierarchy provides a structured taxonomy of object categories, offering a valuable lens through which to analyze the representations learned by deep vision models. In…

cs.AI2025

FastRM: An efficient and automatic explainability framework for multimodal generative models

Gabriela Ben-Melech Stan, Estelle Aflalo, Man Luo +5

Large Vision Language Models (LVLMs) have demonstrated remarkable reasoning capabilities over textual and visual inputs. However, these models remain prone to generating misinforma…