5 papers
LaMI: Augmenting Large Language Models via Late Multi-Image Fusion
Guy Yariv, Idan Schwartz, Yossi Adi +1
Commonsense reasoning often requires both textual and visual knowledge, yet Large Language Models (LLMs) trained solely on text lack visual grounding (e.g., "what color is an emper…
TempoControl: Temporal Attention Guidance for Text-to-Video Models
Shira Schiber, Ofir Lindenbaum, Idan Schwartz
Recent advances in generative video models have enabled the creation of high-quality videos based on natural language prompts. However, these models frequently lack fine-grained te…
Detection-Driven Object Count Optimization for Text-to-Image Diffusion Models
Oz Zafar, Yuval Cohen, Lior Wolf +1
Accurately controlling object count in text-to-image generation remains a key challenge. Supervised methods often fail, as training data rarely covers all count variations. Methods…
Single Image Iterative Subject-driven Generation and Editing
Yair Shpitzer, Gal Chechik, Idan Schwartz
Personalizing image generation and editing is particularly challenging when we only have a few images of the subject, or even a single image. A common approach to personalization i…
Discriminative Class Tokens for Text-to-Image Diffusion Models
Idan Schwartz, Vésteinn Snæbjarnarson, Hila Chefer +4
Recent advances in text-to-image diffusion models have enabled the generation of diverse and high-quality images. While impressive, the images often fall short of depicting subtle…