6 citations · 7 across the 3 of their papers we have counts for
8 papers · 1 filter
Robust Cross-Modal Representation Learning with Progressive Self-Distillation
Alex Andonian, Shixing Chen, Raffay Hamid
The learning objective of vision-language approach of CLIP does not effectively account for the noisy many-to-many correspondences found in web-harvested image captioning datasets,…
The Algonauts Project 2021 Challenge: How the Human Brain Makes Sense of a World in Motion
R. M. Cichy, K. Dwivedi, B. Lahner +8
The sciences of natural and artificial intelligence are fundamentally connected. Brain-inspired human-engineered AI are now the standard for predicting human brain responses during…
VA-RED: Video Adaptive Redundancy Reduction
Bowen Pan, Rameswar Panda, Camilo Fosco +6
Performing inference on deep learning models for videos remains a challenge due to the large amount of computational resources required to achieve robust recognition. An inherent p…
We Have So Much In Common: Modeling Semantic Relational Set Abstractions in Videos
Alex Andonian, Camilo Fosco, Mathew Monfort +4
Identifying common patterns among events is a key ability in human and machine perception, as it underlies intelligent decision making. We propose an approach for learning semantic…
GANalyze: Toward Visual Definitions of Cognitive Image Properties
Lore Goetschalckx, Alex Andonian, Aude Oliva +1
We introduce a framework that uses Generative Adversarial Networks (GANs) to study cognitive properties like memorability, aesthetics, and emotional valence. These attributes are o…
Cross-view Semantic Segmentation for Sensing Surroundings
Bowen Pan, Jiankai Sun, Ho Yin Tiga Leung +2
Sensing surroundings plays a crucial role in human spatial perception, as it extracts the spatial configuration of objects as well as the free space from the observations. To facil…