33 papers
UniProbe: A Learnable Token-Level Hallucination Detector for Large VLMs using Multi-Structural Internal Representations
Dvir Samuel, Guy Bar-Shalom, Fabrizio Frasca +4
Large Vision-Language Models (LVLMs) achieve impressive visual reasoning and dialogue capabilities, yet frequently hallucinate content unsupported by the visual input. Effective mi…
Spanning the Visual Analogy Space with a Weight Basis of LoRAs
Hila Manor, Rinon Gal, Haggai Maron +2
Visual analogy learning enables image editing via demonstration rather than textual description, allowing users to specify complex transformations difficult to articulate in words.…
Fast Autoregressive Video Diffusion and World Models with Temporal Cache Compression and Sparse Attention
Dvir Samuel, Issar Tzachor, Matan Levy +3
Autoregressive video diffusion models enable streaming generation, opening the door to long-form synthesis, video world models, and interactive neural game engines. However, their…
Bootstrap Your Generator: Unpaired Visual Editing with Flow Matching
Yoad Tewel, Yuval Atzmon, Gal Chechik +1
Modern generative models possess a deep understanding of visual content, yet training them for image editing typically requires massive datasets of paired examples. This limits sca…
Expert Merging in Sparse Mixture of Experts with Nash Bargaining
Dung V. Nguyen, Anh T. Nguyen, Minh H. Nguyen +6
Existing expert merging strategies for Sparse Mixture of Experts (SMoE) typically rely on input-dependent or input-independent averaging of expert parameters, but often lack a prin…
Fast 4D Mesh Generation by Spatio-Temporal Attention Chains
Dvir Samuel, Yuval Atzmon, Gal Chechik +1
4D mesh generation has recently emerged as a powerful paradigm for recovering dynamic 3D structure from videos, but existing methods remain slow, computationally expensive, and dif…