26 citations · 49 across the 11 of their papers we have counts for
Showing cs.CVShow all
2 papers · 1 filter
cs.CV2026
Representations of Text and Images Align From Layer One
Evžen Wybitul, Javier Rando, Florian Tramèr +1
We show that for a variety of concepts in adapter-based vision-language models, the representations of their images and their text descriptions are meaningfully aligned from the ve…
cs.CV2024★ 3 cited
Llama Guard 3 Vision: Safeguarding Human-AI Image Understanding Conversations
Jianfeng Chi, Ujjwal Karn, Hongyuan Zhan +7
We introduce Llama Guard 3 Vision, a multimodal LLM-based safeguard for human-AI conversations that involves image understanding: it can be used to safeguard content for both multi…