4 papers · 1 filter
Compositional Video Generation via Inference-Time Guidance
Ariel Shaulov, Eitan Shaar, Amit Edenzon +2
Text-to-video diffusion models generate realistic videos, but often fail on prompts requiring fine-grained compositional understanding, such as relations between entities, attribut…
Latent Transfer Attack: Adversarial Examples via Generative Latent Spaces
Eitan Shaar, Ariel Shaulov, Yalcin Tur +2
Adversarial attacks are a central tool for probing the robustness of modern vision models, yet most methods optimize perturbations directly in pixel space under or $\…
TokenTrim: Inference-Time Token Pruning for Autoregressive Long Video Generation
Ariel Shaulov, Eitan Shaar, Amit Edenzon +1
Auto-regressive video generation enables long video synthesis by iteratively conditioning each new batch of frames on previously generated content. However, recent work has shown t…
Adapting to the Unknown: Training-Free Audio-Visual Event Perception with Dynamic Thresholds
Eitan Shaar, Ariel Shaulov, Gal Chechik +1
In the domain of audio-visual event perception, which focuses on the temporal localization and classification of events across distinct modalities (audio and visual), existing appr…