112 citations · 231 across the 17 of their papers we have counts for
Showing 2022Show all
2 papers · 1 filter
cs.CV2022★ 20 cited
What do Vision Transformers Learn? A Visual Exploration
Amin Ghiasi, Hamid Kazemi, Eitan Borgnia +5
Vision transformers (ViTs) are quickly becoming the de-facto architecture for computer vision, yet we understand very little about why they work and what they learn. While existing…
cs.CV2022★ 112 cited
Test-Time Prompt Tuning for Zero-Shot Generalization in Vision-Language Models
Manli Shu, Weili Nie, De-An Huang +4
Pre-trained vision-language models (e.g., CLIP) have shown promising zero-shot generalization in many downstream tasks with properly designed text prompts. Instead of relying on ha…