5 papers · 1 filter
PREGEN: Uncovering Latent Thoughts in Composed Video Retrieval
Gabriele Serussi, David Vainshtein, Jonathan Kouchly +2
Composed Video Retrieval (CoVR) aims to retrieve a video based on a query video and a modifying text. Current CoVR methods fail to fully exploit modern Vision-Language Models (VLMs…
Towards General Modality Translation with Contrastive and Predictive Latent Diffusion Bridge
Nimrod Berman, Omkar Joglekar, Eitan Kosman +2
Recent advances in generative modeling have positioned diffusion models as state-of-the-art tools for sampling from complex data distributions. While these models have shown remark…
Robot Instance Segmentation with Few Annotations for Grasping
Moshe Kimhi, David Vainshtein, Chaim Baskin +1
The ability of robots to manipulate objects relies heavily on their aptitude for visual perception. In domains characterized by cluttered scenes and high object variability, most m…
Radar Spectra-Language Model for Automotive Scene Parsing
Mariia Pushkareva, Yuri Feldman, Csaba Domokos +2
Radar sensors are low cost, long-range, and weather-resilient. Therefore, they are widely used for driver assistance functions, and are expected to be crucial for the success of au…
ISCUTE: Instance Segmentation of Cables Using Text Embedding
Shir Kozlovsky, Omkar Joglekar, Dotan Di Castro
In the field of robotics and automation, conventional object recognition and instance segmentation methods face a formidable challenge when it comes to perceiving Deformable Linear…