activity
20242026
collaborators
Showing cs.CVShow all

5 papers · 1 filter

cs.CV2026

TokenTrim: Inference-Time Token Pruning for Autoregressive Long Video Generation

Ariel Shaulov, Eitan Shaar, Amit Edenzon +1

Auto-regressive video generation enables long video synthesis by iteratively conditioning each new batch of frames on previously generated content. However, recent work has shown t…

cs.CV2025

Detection-Driven Object Count Optimization for Text-to-Image Diffusion Models

Oz Zafar, Yuval Cohen, Lior Wolf +1

Accurately controlling object count in text-to-image generation remains a key challenge. Supervised methods often fail, as training data rarely covers all count variations. Methods…

cs.CV2025

Adapting to the Unknown: Training-Free Audio-Visual Event Perception with Dynamic Thresholds

Eitan Shaar, Ariel Shaulov, Gal Chechik +1

In the domain of audio-visual event perception, which focuses on the temporal localization and classification of events across distinct modalities (audio and visual), existing appr…

cs.CV2024

Add-it: Training-Free Object Insertion in Images With Pretrained Diffusion Models

Yoad Tewel, Rinon Gal, Dvir Samuel +3

Adding Object into images based on text instructions is a challenging task in semantic image editing, requiring a balance between preserving the original scene and seamlessly integ…

cs.CV2024

Training-Free Consistent Text-to-Image Generation

Yoad Tewel, Omri Kaduri, Rinon Gal +4

Text-to-image models offer a new level of creative flexibility by allowing users to guide the image generation process through natural language. However, using these models to cons…