works on

From the 1 of 7 linked papers with an AI index.

activity
20242026
collaborators
Showing cs.CVShow all

5 papers · 1 filter

cs.CV2026

Multimodality as Supervision: Self-Supervised Specialization to the Test Environment via Multimodality

Kunal Pratap Singh, Ali Garjani, Rishubh Singh +6

The paper introduces Test-Space Training, a self‑supervised approach that collects multimodal sensor data directly in a target test environment and uses cross‑modal learning to pre…

cs.CV2026

How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks

Rahul Ramachandran, Ali Garjani, Roman Bachmann +3

Multimodal foundation models (MFMs), such as GPT-4o, have recently made remarkable progress. However, their detailed visual understanding beyond question answering remains unclear.…

cs.CV2026

VideoFlexTok: Flexible-Length Coarse-to-Fine Video Tokenization

Andrei Atanov, Jesse Allardice, Roman Bachmann +6

Visual tokenizers map high-dimensional raw pixels into a compressed representation for downstream modeling. Beyond compression, tokenizers dictate what information is preserved and…

cs.CV2025

Controlled Training Data Generation with Diffusion Models

Teresa Yeo, Andrei Atanov, Harold Benoit +4

We present a method to control a text-to-image generative model to produce training data useful for supervised learning. Unlike previous works that employ an open-loop approach and…

cs.CV2024

Solving Vision Tasks with Simple Photoreceptors Instead of Cameras

Andrei Atanov, Jiawei Fu, Rishubh Singh +3

A de facto standard in solving computer vision problems is to use a common high-resolution camera and choose its placement on an agent (i.e., position and orientation) based on hum…