4 papers
HiMu: Hierarchical Multimodal Frame Selection for Long Video Question Answering
Dan Ben-Ami, Gabriele Serussi, Kobi Cohen +1
Long-form video question answering requires reasoning over extended temporal contexts, making frame selection a critical bottleneck for multi-modal large language models (MLLMs) bo…
Structured Diffusion Bridges: Inductive Bias for Denoising Diffusion Bridges
Eitan Kosman, Gabriele Serussi, Chaim Baskin
Modality translation is inherently under-constrained, as multiple cross-modal mappings may yield the same marginals. Recent work has shown that diffusion bridges are effective for…
HERBench: A Benchmark for Multi-Evidence Integration in Video Question Answering
Dan Ben-Ami, Gabriele Serussi, Kobi Cohen +1
Video Large Language Models (Video-LLMs) are improving rapidly, yet current Video Question Answering (VideoQA) benchmarks often admit single-cue shortcuts, under-testing reasoning…
PREGEN: Uncovering Latent Thoughts in Composed Video Retrieval
Gabriele Serussi, David Vainshtein, Jonathan Kouchly +2
Composed Video Retrieval (CoVR) aims to retrieve a video based on a query video and a modifying text. Current CoVR methods fail to fully exploit modern Vision-Language Models (VLMs…