21 papers
Instruction-Based Video Editing by Repurposing an Image Editing Model
Yunpeng Bai, Yossi Gandelsman, Michaël Gharbi +1
Instruction-based video editing is commonly built on video-pretrained generative backbones: a video diffusion transformer is adapted, at considerable cost, to condition on a source…
Learning Sampling Parameters for Diffusion Models
Arisrei Lim, Yossi Gandelsman
Text-to-image diffusion models expose many inference-time sampling parameters, including prompts, negative prompts, classifier-free guidance scales, and noise schedules. These para…
Test-Time Training for Modality Order Consistency in Vision-Language Models
Aditi Gupta, Yossi Gandelsman
We find that vision-language models are sensitive to a specific semantically irrelevant change: the order in which the image and question are presented. Across three models and thr…
DuSPiT: Dual-Branch Sub-Patch Pixel Diffusion Transformer
Yunpeng Bai, Yossi Gandelsman, Michaël Gharbi
Diffusion Transformers achieve strong image generation performance, but most operate in compressed latent spaces. Pixel-space diffusion avoids this information loss, yet existing a…
Neuron Populations Exhibit Divergent Selectivity with Scale
Amil Dravid, Yasaman Bahri, Alexei A. Efros +1
We investigate whether neuron populations within neural networks evolve predictably with scale, extending scaling laws beyond macroscopic observables such as loss. To probe this qu…
Jailbreaking Vision-Language Models Through the Visual Modality
Aharon Azulay, Jan DubiÅski, Zhuoyun Li +2
The visual modality of vision-language models (VLMs) is an underexplored attack surface for bypassing safety alignment. We introduce four jailbreak attacks exploiting the vision co…