activity
20202025
most citedTest-Time Training with Masked Autoencoders

36 citations · 36 across the 4 of their papers we have counts for

collaborators

7 papers

cs.CV2025

LLMs can see and hear without any training

Kumar Ashutosh, Yossi Gandelsman, Xinlei Chen +2

We present MILS: Multimodal Iterative LLM Solver, a surprisingly simple, training-free approach, to imbue multimodal capabilities into your favorite LLM. Leveraging their innate ab…

cs.CV2025

An Empirical Study of Autoregressive Pre-training from Videos

Jathushan Rajasegaran, Ilija Radosavovic, Rahul Ravishankar +3

We empirically study autoregressive pre-training from videos. To perform our study, we construct a series of autoregressive video models, called Toto. We treat videos as sequences…

cs.CV2024

Quantifying and Enabling the Interpretability of CLIP-like Models

Avinash Madasu, Yossi Gandelsman, Vasudev Lal +1

CLIP is one of the most popular foundational models and is heavily used for many vision-language tasks. However, little is known about the inner workings of CLIP. To bridge this ga…

cs.CV202236 cited

Test-Time Training with Masked Autoencoders

Yossi Gandelsman, Yu Sun, Xinlei Chen +1

Test-time training adapts to a new test distribution on the fly by optimizing a model for each test input using self-supervision. In this paper, we use masked autoencoders for this…

cs.CV2021

Deep Saliency Prior for Reducing Visual Distraction

Kfir Aberman, Junfeng He, Yossi Gandelsman +5

Using only a model that was trained to predict where people look at images, and no additional training data, we can produce a range of powerful editing effects for reducing distrac…

cs.CV2021

Explaining in Style: Training a GAN to explain a classifier in StyleSpace

Oran Lang, Yossi Gandelsman, Michal Yarom +8

Image classification models can depend on multiple different semantic attributes of the image. An explanation of the decision of the classifier needs to both discover and visualize…