3 papers
cs.LG2024
Exploring Curriculum Learning for Vision-Language Tasks: A Study on Small-Scale Multimodal Training
Rohan Saha, Abrar Fahim, Alona Fyshe +1
For specialized domains, there is often not a wealth of data with which to train large machine learning models. In such limited data / compute settings, various methods exist aimin…
cs.LG2024
Finding Shared Decodable Concepts and their Negations in the Brain
Cory Efird, Alex Murphy, Joel Zylberberg +1
Prior work has offered evidence for functional localization in the brain; different anatomical regions preferentially activate for certain types of visual input. For example, the f…
cs.CV2024
It's Not a Modality Gap: Characterizing and Addressing the Contrastive Gap
Abrar Fahim, Alex Murphy, Alona Fyshe
Multi-modal contrastive models such as CLIP achieve state-of-the-art performance in zero-shot classification by embedding input images and texts on a joint representational space.…