activity
20242026
collaborators

6 papers

cs.CV2026

Do multimodal models imagine electric sheep?

Santhosh Kumar Ramakrishnan, Carl Vondrick, Raja Giryes +2

Yes. We find that large multimodal models develop mental imagery when solving spatial puzzles, and they do imagine sheep when solving sheep puzzles. We fine-tune a Qwen3.5 VLM to s…

cs.CV2026

Video Analysis and Generation via a Semantic Progress Function

Gal Metzer, Sagi Polaczek, Ali Mahdavi-Amiri +2

Transformations produced by image and video generation models often evolve in a highly non-linear manner: long stretches where the content barely changes are followed by sudden, ab…

cs.CV2026

Anatomical Token Uncertainty for Transformer-Guided Active MRI Acquisition

Lev Ayzenberg, Shady Abu-Hussein, Raja Giryes +1

Full data acquisition in MRI is inherently slow, which limits clinical throughput and increases patient discomfort. Compressed Sensing MRI (CS-MRI) seeks to accelerate acquisition…

cs.SD2026

ID-LoRA: Identity-Driven Audio-Video Personalization with In-Context LoRA

Aviad Dahan, Moran Yanuka, Noa Kraicer +2

Existing video personalization methods preserve visual likeness but treat video and audio separately. Without access to the visual scene, audio models cannot synchronize sounds wit…

eess.IV2025

X-ray2CTPA: Leveraging Diffusion Models to Enhance Pulmonary Embolism Classification

Noa Cahan, Eyal Klang, Galit Aviram +4

Chest X-rays or chest radiography (CXR), commonly used for medical diagnostics, typically enables limited imaging compared to computed tomography (CT) scans, which offer more detai…

cs.LG2024

On the Relation Between Linear Diffusion and Power Iteration

Dana Weitzner, Mauricio Delbracio, Peyman Milanfar +1

Recently, diffusion models have gained popularity due to their impressive generative abilities. These models learn the implicit distribution given by the training dataset, and samp…