5 papers
Multimodality as Supervision: Self-Supervised Specialization to the Test Environment via Multimodality
Kunal Pratap Singh, Ali Garjani, Rishubh Singh +6
The paper introduces Test-Space Training, a self‑supervised approach that collects multimodal sensor data directly in a target test environment and uses cross‑modal learning to pre…
How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks
Rahul Ramachandran, Ali Garjani, Roman Bachmann +3
Multimodal foundation models (MFMs), such as GPT-4o, have recently made remarkable progress. However, their detailed visual understanding beyond question answering remains unclear.…
VideoFlexTok: Flexible-Length Coarse-to-Fine Video Tokenization
Andrei Atanov, Jesse Allardice, Roman Bachmann +6
Visual tokenizers map high-dimensional raw pixels into a compressed representation for downstream modeling. Beyond compression, tokenizers dictate what information is preserved and…
Controlled Training Data Generation with Diffusion Models
Teresa Yeo, Andrei Atanov, Harold Benoit +4
We present a method to control a text-to-image generative model to produce training data useful for supervised learning. Unlike previous works that employ an open-loop approach and…
Large (Vision) Language Models are Unsupervised In-Context Learners
Artyom Gadetsky, Andrei Atanov, Yulun Jiang +4
Recent advances in large language and vision-language models have enabled zero-shot inference, allowing models to solve new tasks without task-specific training. Various adaptation…