3 citations · 4 across the 24 of their papers we have counts for
3 papers · 1 filter
Representations of Text and Images Align From Layer One
Evžen Wybitul, Javier Rando, Florian Tramèr +1
We show that for a variety of concepts in adapter-based vision-language models, the representations of their images and their text descriptions are meaningfully aligned from the ve…
Modal Aphasia: Can Unified Multimodal Models Describe Images From Memory?
Michael Aerni, Joshua Swanson, Kristina Nikolić +1
We present modal aphasia, a systematic dissociation in which current unified multimodal models accurately memorize concepts visually but fail to articulate them in writing, despite…
Extracting Training Data from Document-Based VQA Models
Francesco Pinto, Nathalie Rauschmayr, Florian Tramèr +2
Vision-Language Models (VLMs) have made remarkable progress in document-based Visual Question Answering (i.e., responding to queries about the contents of an input document provide…