6 papers
Large (Vision) Language Models are Unsupervised In-Context Learners
Artyom Gadetsky, Andrei Atanov, Yulun Jiang +4
Recent advances in large language and vision-language models have enabled zero-shot inference, allowing models to solve new tasks without task-specific training. Various adaptation…
Solving Vision Tasks with Simple Photoreceptors Instead of Cameras
Andrei Atanov, Jiawei Fu, Rishubh Singh +3
A de facto standard in solving computer vision problems is to use a common high-resolution camera and choose its placement on an agent (i.e., position and orientation) based on hum…
4M-21: An Any-to-Any Vision Model for Tens of Tasks and Modalities
Roman Bachmann, Oğuzhan Fatih Kar, David Mizrahi +6
Current multimodal and multitask foundation models like 4M or UnifiedIO show promising results, but in practice their out-of-the-box abilities to accept diverse inputs and perform…
BRAVE: Broadening the visual encoding of vision-language models
Oğuzhan Fatih Kar, Alessio Tonioni, Petra Poklukar +3
Vision-language models (VLMs) are typically composed of a vision encoder, e.g. CLIP, and a language model (LM) that interprets the encoded features to solve downstream tasks. Despi…
Rapid Network Adaptation: Learning to Adapt Neural Networks Using Test-Time Feedback
Teresa Yeo, Oğuzhan Fatih Kar, Zahra Sodagar +1
We propose a method for adapting neural networks to distribution shifts at test-time. In contrast to training-time robustness mechanisms that attempt to anticipate and counter the…
Modality-invariant Visual Odometry for Embodied Vision
Marius Memmel, Roman Bachmann, Amir Zamir
Effectively localizing an agent in a realistic, noisy setting is crucial for many embodied vision tasks. Visual Odometry (VO) is a practical substitute for unreliable GPS and compa…