4.4k citations · 4.6k across the 13 of their papers we have counts for
17 papers · 1 filter
4M: Massively Multimodal Masked Modeling
David Mizrahi, Roman Bachmann, Oğuzhan Fatih Kar +4
Current machine learning models for vision are often highly specialized and limited to a single modality and task. In contrast, recent large language models exhibit a wide range of…
Rapid Network Adaptation: Learning to Adapt Neural Networks Using Test-Time Feedback
Teresa Yeo, Oğuzhan Fatih Kar, Zahra Sodagar +1
We propose a method for adapting neural networks to distribution shifts at test-time. In contrast to training-time robustness mechanisms that attempt to anticipate and counter the…
Modality-invariant Visual Odometry for Embodied Vision
Marius Memmel, Roman Bachmann, Amir Zamir
Effectively localizing an agent in a realistic, noisy setting is crucial for many embodied vision tasks. Visual Odometry (VO) is a practical substitute for unreliable GPS and compa…
3D Common Corruptions and Data Augmentation
Oğuzhan Fatih Kar, Teresa Yeo, Andrei Atanov +1
We introduce a set of image transformations that can be used as corruptions to evaluate the robustness of models as well as data augmentation mechanisms for training neural network…
MultiMAE: Multi-modal Multi-task Masked Autoencoders
Roman Bachmann, David Mizrahi, Andrei Atanov +1
We propose a pre-training strategy called Multi-modal Multi-task Masked Autoencoders (MultiMAE). It differs from standard Masked Autoencoding in two key aspects: I) it can optional…
Simple Control Baselines for Evaluating Transfer Learning
Andrei Atanov, Shijian Xu, Onur Beker +2
Transfer learning has witnessed remarkable progress in recent years, for example, with the introduction of augmentation-based contrastive self-supervised learning methods. While a…