activity
20122023
most citedUCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild

4.4k citations · 4.6k across the 13 of their papers we have counts for

collaborators
Showing cs.CVShow all

17 papers · 1 filter

cs.CV2023

4M: Massively Multimodal Masked Modeling

David Mizrahi, Roman Bachmann, Oğuzhan Fatih Kar +4

Current machine learning models for vision are often highly specialized and limited to a single modality and task. In contrast, recent large language models exhibit a wide range of…

cs.CV2023

Rapid Network Adaptation: Learning to Adapt Neural Networks Using Test-Time Feedback

Teresa Yeo, Oğuzhan Fatih Kar, Zahra Sodagar +1

We propose a method for adapting neural networks to distribution shifts at test-time. In contrast to training-time robustness mechanisms that attempt to anticipate and counter the…

cs.CV2023

Modality-invariant Visual Odometry for Embodied Vision

Marius Memmel, Roman Bachmann, Amir Zamir

Effectively localizing an agent in a realistic, noisy setting is crucial for many embodied vision tasks. Visual Odometry (VO) is a practical substitute for unreliable GPS and compa…

cs.CV20224 cited

3D Common Corruptions and Data Augmentation

Oğuzhan Fatih Kar, Teresa Yeo, Andrei Atanov +1

We introduce a set of image transformations that can be used as corruptions to evaluate the robustness of models as well as data augmentation mechanisms for training neural network…

cs.CV202210 cited

MultiMAE: Multi-modal Multi-task Masked Autoencoders

Roman Bachmann, David Mizrahi, Andrei Atanov +1

We propose a pre-training strategy called Multi-modal Multi-task Masked Autoencoders (MultiMAE). It differs from standard Masked Autoencoding in two key aspects: I) it can optional…

cs.CV20221 cited

Simple Control Baselines for Evaluating Transfer Learning

Andrei Atanov, Shijian Xu, Onur Beker +2

Transfer learning has witnessed remarkable progress in recent years, for example, with the introduction of augmentation-based contrastive self-supervised learning methods. While a…