6 citations · 8 across the 18 of their papers we have counts for
6 papers · 1 filter
Unsupervised Universal Image Segmentation
Dantong Niu, Xudong Wang, Xinyang Han +3
Several unsupervised image segmentation approaches have been proposed which eliminate the need for dense manually-annotated segmentation masks; current models separately handle eit…
Recursive Visual Programming
Jiaxin Ge, Sanjay Subramanian, Baifeng Shi +2
Visual Programming (VP) has emerged as a powerful framework for Visual Question Answering (VQA). By generating and executing bespoke code for each question, these methods demonstra…
Object-based (yet Class-agnostic) Video Domain Adaptation
Dantong Niu, Amir Bar, Roei Herzig +2
Existing video-based action recognition systems typically require dense annotation and struggle in environments when there is significant distribution shift relative to the trainin…
Compositional Chain-of-Thought Prompting for Large Multimodal Models
Chancharik Mitra, Brandon Huang, Trevor Darrell +1
The combination of strong visual backbones and Large Language Model (LLM) reasoning has led to Large Multimodal Models (LMMs) becoming the current standard for a wide range of visi…
Dense and Aligned Captions (DAC) Promote Compositional Reasoning in VL Models
Sivan Doveh, Assaf Arbelle, Sivan Harary +9
Vision and Language (VL) models offer an effective method for aligning representation spaces of images and text, leading to numerous applications such as cross-modal retrieval, vis…
Incorporating Structured Representations into Pretrained Vision & Language Models Using Scene Graphs
Roei Herzig, Alon Mendelson, Leonid Karlinsky +4
Vision and language models (VLMs) have demonstrated remarkable zero-shot (ZS) performance in a variety of tasks. However, recent works have shown that even the best VLMs struggle t…