3 papers
cs.CV2024
Nearest Neighbor Normalization Improves Multimodal Retrieval
Neil Chowdhury, Franklin Wang, Sumedh Shenoy +3
Multimodal models leverage large-scale pre-training to achieve strong but still imperfect performance on tasks such as image captioning, visual question answering, and cross-modal…
cs.CV2024
Automatic Discovery of Visual Circuits
Achyuta Rajaram, Neil Chowdhury, Antonio Torralba +2
To date, most discoveries of network subcomponents that implement human-interpretable computations in deep vision models have involved close study of single units and large amounts…
cs.CV2023
Multimodal Neurons in Pretrained Text-Only Transformers
Sarah Schwettmann, Neil Chowdhury, Samuel Klein +2
Language models demonstrate remarkable capacity to generalize representations learned in one modality to downstream tasks in other modalities. Can we trace this ability to individu…