6 papers
The Universal Normal Embedding
Chen Tasker, Roy Betser, Eyal Gofer +2
Generative models and vision encoders have largely advanced on separate tracks, optimized for different goals and grounded in different mathematical principles. Yet, they share a f…
Training-free Detection of Generated Videos via Spatial-Temporal Likelihoods
Omer Ben Hayun, Roy Betser, Meir Yossef Levi +2
Following major advances in text and image generation, the video domain has surged, producing highly realistic and controllable sequences. Along with this progress, these models al…
Make it SING: Analyzing Semantic Invariants in Classifiers
Harel Yadid, Meir Yossef Levi, Roy Betser +1
All classifiers, including state-of-the-art vision models, possess invariants, partially rooted in the geometry of their linear mappings. These invariants, which reside in the null…
SCoCCA: Multi-modal Sparse Concept Decomposition via Canonical Correlation Analysis
Ehud Gordon, Meir Yossef Levi, Guy Gilboa
Interpreting the internal reasoning of vision-language models is essential for deploying AI in safety-critical domains. Concept-based explainability provides a human-aligned lens b…
Whitened CLIP as a Likelihood Surrogate of Images and Captions
Roy Betser, Meir Yossef Levi, Guy Gilboa
Likelihood approximations for images are not trivial to compute and can be useful in many applications. We examine the use of Contrastive Language-Image Pre-training (CLIP) to asse…
The Double-Ellipsoid Geometry of CLIP
Meir Yossef Levi, Guy Gilboa
Contrastive Language-Image Pre-Training (CLIP) is highly instrumental in machine learning applications within a large variety of domains. We investigate the geometry of this embedd…