6 papers
Modular Diffusion Models for Structured Visual Recognition
Siddhesh Khandelwal, Björn Ommer, Leonid Sigal
Traditional supervised methods for structured visual recognition tasks -- such as object detection, segmentation, and scene graph generation -- often produce deterministic, fixed o…
The ART of Composition: Attention-Regularized Training for Compositional Visual Grounding
Jiayun Luo, Mir Rayat Imtiaz Hossain, Pritam Sarkar +2
Vision-Language Models (VLMs) have achieved strong performance on implicit and explicit visual grounding and related tasks. However, such abilities are generally tested on simple,…
Do LLMs Benefit from User and Item Embeddings in Recommendation Tasks?
Mir Rayat Imtiaz Hossain, Leo Feng, Leonid Sigal +1
Large Language Models (LLMs) have emerged as promising recommendation systems, offering novel ways to model user preferences through generative approaches. However, many existing m…
ChartGaze: Enhancing Chart Understanding in LVLMs with Eye-Tracking Guided Attention Refinement
Ali Salamatian, Amirhossein Abaskohi, Wan-Cyuan Fan +3
Charts are a crucial visual medium for communicating and representing information. While Large Vision-Language Models (LVLMs) have made progress on chart question answering (CQA),…
Black Swan: Abductive and Defeasible Video Reasoning in Unpredictable Events
Aditya Chinchure, Sahithya Ravi, Raymond Ng +3
The commonsense reasoning capabilities of vision-language models (VLMs), especially in abductive reasoning and defeasible reasoning, remain poorly understood. Most benchmarks focus…
The Power of One: A Single Example is All it Takes for Segmentation in VLMs
Mir Rayat Imtiaz Hossain, Mennatullah Siam, Leonid Sigal +1
Large-scale vision-language models (VLMs), trained on extensive datasets of image-text pairs, exhibit strong multimodal understanding capabilities by implicitly learning associatio…