6 papers
Token-to-Token Alignment of Text Embeddings for Semantic Blending
Saar Huberman, Ron Mokady, Or Patashnik +1
In modern generative models, images are specified and controlled through text prompts. In practice, images are generated from sequences of tokens derived from these prompts. Howeve…
Image Generation from Contextually-Contradictory Prompts
Saar Huberman, Or Patashnik, Omer Dahary +2
Text-to-image diffusion models excel at generating high-quality, diverse images from natural language prompts. However, they often fail to produce semantically accurate results whe…
Deep Accurate Solver for the Geodesic Problem
Saar Huberman, Amit Bracha, Ron Kimmel
A common approach to compute distances on continuous surfaces is by considering a discretized polygonal mesh approximating the surface and estimating distances on the polygon. We s…
BBQ-to-Image: Numeric Bounding Box and Qolor Control in Large-Scale Text-to-Image Models
Eliran Kachlon, Alexander Visheratin, Nimrod Sarid +6
Text-to-image models have rapidly advanced in realism and controllability, with recent approaches leveraging long, detailed captions to support fine-grained generation. However, a…
SemanticMoments: Training-Free Motion Similarity via Third Moment Features
Saar Huberman, Kfir Goldberg, Or Patashnik +2
Retrieving videos based on semantic motion is a fundamental, yet unsolved, problem. Existing video representation approaches overly rely on static appearance and scene context rath…
Generating an Image From 1,000 Words: Enhancing Text-to-Image With Structured Captions
Eyal Gutflaish, Eliran Kachlon, Hezi Zisman +8
Text-to-image models have rapidly evolved from casual creative tools to professional-grade systems, achieving unprecedented levels of image quality and realism. Yet, most models ar…