Showing cs.CVShow all
2 papers · 1 filter
cs.CV2025
Vision Transformers Don't Need Trained Registers
Nick Jiang, Amil Dravid, Alexei Efros +1
We investigate the mechanism underlying a previously identified phenomenon in Vision Transformers - the emergence of high-norm tokens that lead to noisy attention maps (Darcet et a…
cs.CV2025
Interpreting and Editing Vision-Language Representations to Mitigate Hallucinations
Nick Jiang, Anish Kachinthaya, Suzie Petryk +1
We investigate the internal representations of vision-language models (VLMs) to address hallucinations, a persistent challenge despite advances in model size and training. We proje…