2 papers
cs.CV2026
Bi-Orthogonal Factor Decomposition for Vision Transformers
Fenil R. Doshi, Thomas Fel, Talia Konkle +1
Self-attention is the central computational primitive of Vision Transformers, yet we lack a principled understanding of what information attention mechanisms exchange between token…
cs.CV2025
Visual Anagrams Reveal Hidden Differences in Holistic Shape Processing Across Vision Models
Fenil R. Doshi, Thomas Fel, Talia Konkle +1
Humans are able to recognize objects based on both local texture cues and the configuration of object parts, yet contemporary vision models primarily harvest local texture cues, yi…