4 papers
Brain-IT-VQA: From Brain Signals to Answers
Roman Beliy, Matias Cosarinsky, Oliver Heinimann +2
Decoding visual content from fMRI signals recorded while a person views images, and specifically answering questions about the seen images, is a long-standing challenge. While sign…
DocReRank: Single-Page Hard Negative Query Generation for Training Multi-Modal RAG Rerankers
Navve Wasserman, Oliver Heinimann, Yuval Golbari +3
Rerankers play a critical role in multimodal Retrieval-Augmented Generation (RAG) by refining ranking of an initial set of retrieved documents. Rerankers are typically trained usin…
KernelFusion: Assumption-Free Blind Super-Resolution via Patch Diffusion
Oliver Heinimann, Assaf Shocher, Tal Zimbalist +1
Traditional super-resolution (SR) methods assume an ``ideal'' downscaling SR-kernel (e.g., bicubic downscaling) between the high-resolution (HR) image and the low-resolution (LR) i…
Don't Judge Before You CLIP: A Unified Approach for Perceptual Tasks
Amit Zalcher, Navve Wasserman, Roman Beliy +2
Visual perceptual tasks aim to predict human judgment of images (e.g., emotions invoked by images, image quality assessment). Unlike objective tasks such as object/scene recognitio…