9 papers
LieCraft: A Multi-Agent Framework for Evaluating Deceptive Capabilities in Language Models
Matthew Lyle Olson, Neale Ratzlaff, Musashi Hinck +5
Large Language Models (LLMs) exhibit impressive general-purpose capabilities but also introduce serious safety risks, particularly the potential for deception as models acquire inc…
Probing the Representational Power of Sparse Autoencoders in Vision Models
Matthew Lyle Olson, Musashi Hinck, Neale Ratzlaff +4
Sparse Autoencoders (SAEs) have emerged as a popular tool for interpreting the hidden states of large language models (LLMs). By learning to reconstruct activations from a sparse b…
DPO Learning with LLMs-Judge Signal for Computer Use Agents
Man Luo, David Cobbley, Xin Su +4
Computer use agents (CUA) are systems that automatically interact with graphical user interfaces (GUIs) to complete tasks. CUA have made significant progress with the advent of lar…
Analyzing Hierarchical Structure in Vision Models with Sparse Autoencoders
Matthew Lyle Olson, Musashi Hinck, Neale Ratzlaff +4
The ImageNet hierarchy provides a structured taxonomy of object categories, offering a valuable lens through which to analyze the representations learned by deep vision models. In…
Steering Large Language Models to Evaluate and Amplify Creativity
Matthew Lyle Olson, Neale Ratzlaff, Musashi Hinck +2
Although capable of generating creative text, Large Language Models (LLMs) are poor judges of what constitutes "creativity". In this work, we show that we can leverage this knowled…
FiVL: A Framework for Improved Vision-Language Alignment through the Lens of Training, Evaluation and Explainability
Estelle Aflalo, Gabriela Ben Melech Stan, Tiep Le +5
Large Vision Language Models (LVLMs) have achieved significant progress in integrating visual and textual inputs for multimodal reasoning. However, a recurring challenge is ensurin…