15 papers
UniProbe: A Learnable Token-Level Hallucination Detector for Large VLMs using Multi-Structural Internal Representations
Dvir Samuel, Guy Bar-Shalom, Fabrizio Frasca +4
Large Vision-Language Models (LVLMs) achieve impressive visual reasoning and dialogue capabilities, yet frequently hallucinate content unsupported by the visual input. Effective mi…
WarpHammer: Densifying Scene Warps with 3D Object Priors for Extreme View Synthesis
Michael Green, Gavriel Habib, Dvir Samuel +4
Projection-conditioned novel view synthesis (NVS) warps an explicit 3D reconstruction of the input view into the target camera and conditions a generator on the warped rendering. T…
Fast Autoregressive Video Diffusion and World Models with Temporal Cache Compression and Sparse Attention
Dvir Samuel, Issar Tzachor, Matan Levy +3
Autoregressive video diffusion models enable streaming generation, opening the door to long-form synthesis, video world models, and interactive neural game engines. However, their…
LiveSVG: Zero-Shot SVG Animation via Video Generation
Matan Levy, Ran Margolin, Bar Cavia +6
We introduce LiveSVG, a zero-shot approach for generating Scalable Vector Graphics (SVG) animations using video diffusion models. Current SVG animation methods struggle with comple…
Fast 4D Mesh Generation by Spatio-Temporal Attention Chains
Dvir Samuel, Yuval Atzmon, Gal Chechik +1
4D mesh generation has recently emerged as a powerful paradigm for recovering dynamic 3D structure from videos, but existing methods remain slow, computationally expensive, and dif…
Retrieval-Augmented Gaussian Avatars: Improving Expression Generalization
Matan Levy, Gavriel Habib, Issar Tzachor +5
Template-free animatable head avatars can achieve high visual fidelity by learning expression-dependent facial deformation directly from a subject's capture, avoiding parametric fa…