collaborators

18 papers

cs.CV2026

Analysis-by-Proxy: Localization Signals in VLMs Operating as Condition Encoders

Yoav Baron, Sara Dorfman, Roni Paiss +2

Vision-Language Models (VLMs) are increasingly utilized as the conditioning backbone for diffusion-based image editing due to their remarkable multimodal reasoning capabilities. Wh…

cs.CV2026

Token-to-Token Alignment of Text Embeddings for Semantic Blending

Saar Huberman, Ron Mokady, Or Patashnik +1

In modern generative models, images are specified and controlled through text prompts. In practice, images are generated from sequences of tokens derived from these prompts. Howeve…

cs.CV2026

Keep The Essentials: Efficient Reference Conditioned Generation via Token Dropping

Rishubh Parihar, Ayush Raina, R. Venkatesh Babu +1

Reference-based diffusion models enable highly controllable image generation by leveraging elements from input images to guide prompt-driven synthesis. However, these models are co…

cs.CV2026

Semantic Browsing: Controllable Diversity for Image Generation

Sara Dorfman, Maya Vishnevsky, Omer Dahary +2

Modern text-to-image models excel in visual fidelity and prompt adherence. However, this strict adherence comes at the cost of diversity: generated samples tend to collapse into a…

cs.SD2026

Reference-Driven Multi-Speaker Audio Scene Generation from In-the-Wild Priors

Michael Finkelson, Daniel Segal, Eitan Richardson +7

Existing multi-speaker dialogue systems bind speakers to utterances through structured supervision: per-turn tags, multi-stream transcriptions, or learnable speaker embeddings. The…

cs.GR2026

JUST-DUB-IT: Video Dubbing via Joint Audio-Visual Diffusion

Anthony Chen, Naomi Ken Korem, Gal Zeevi +6

Audio-Visual Foundation Models, which are pretrained to jointly generate sound and visual content, have recently shown an unprecedented ability to model multi-modal generation and…