8 papers
WarpHammer: Densifying Scene Warps with 3D Object Priors for Extreme View Synthesis
Michael Green, Gavriel Habib, Dvir Samuel +4
Projection-conditioned novel view synthesis (NVS) warps an explicit 3D reconstruction of the input view into the target camera and conditions a generator on the warped rendering. T…
Fast Autoregressive Video Diffusion and World Models with Temporal Cache Compression and Sparse Attention
Dvir Samuel, Issar Tzachor, Matan Levy +3
Autoregressive video diffusion models enable streaming generation, opening the door to long-form synthesis, video world models, and interactive neural game engines. However, their…
VidMsg: A Benchmark for Implicit Message Inference in Short Videos
Issar Tzachor, Michael Green, Rami Ben-Ari
Understanding short online videos involves more than identifying visible objects and actions; video makers often include an underlying message or purpose in the clip. We introduce…
Retrieval-Augmented Gaussian Avatars: Improving Expression Generalization
Matan Levy, Gavriel Habib, Issar Tzachor +5
Template-free animatable head avatars can achieve high visual fidelity by learning expression-dependent facial deformation directly from a subject's capture, avoiding parametric fa…
VidVec: Unlocking Video MLLM Embeddings for Video-Text Retrieval
Issar Tzachor, Dvir Samuel, Rami Ben-Ari
Recent studies have adapted generative Multimodal Large Language Models (MLLMs) into embedding extractors for vision tasks, typically through fine-tuning to produce universal repre…
Find your Needle: Small Object Image Retrieval via Multi-Object Attention Optimization
Michael Green, Matan Levy, Issar Tzachor +3
We address the challenge of Small Object Image Retrieval (SoIR), where the goal is to retrieve images containing a specific small object, in a cluttered scene. The key challenge in…