collaborators

6 papers

cs.CV2026

Modular Diffusion Models for Structured Visual Recognition

Siddhesh Khandelwal, Björn Ommer, Leonid Sigal

Traditional supervised methods for structured visual recognition tasks -- such as object detection, segmentation, and scene graph generation -- often produce deterministic, fixed o…

cs.CV2026

The ART of Composition: Attention-Regularized Training for Compositional Visual Grounding

Jiayun Luo, Mir Rayat Imtiaz Hossain, Pritam Sarkar +2

Vision-Language Models (VLMs) have achieved strong performance on implicit and explicit visual grounding and related tasks. However, such abilities are generally tested on simple,…

cs.LG2026

Do LLMs Benefit from User and Item Embeddings in Recommendation Tasks?

Mir Rayat Imtiaz Hossain, Leo Feng, Leonid Sigal +1

Large Language Models (LLMs) have emerged as promising recommendation systems, offering novel ways to model user preferences through generative approaches. However, many existing m…

cs.CL2025

ChartGaze: Enhancing Chart Understanding in LVLMs with Eye-Tracking Guided Attention Refinement

Ali Salamatian, Amirhossein Abaskohi, Wan-Cyuan Fan +3

Charts are a crucial visual medium for communicating and representing information. While Large Vision-Language Models (LVLMs) have made progress on chart question answering (CQA),…

cs.CV2025

Black Swan: Abductive and Defeasible Video Reasoning in Unpredictable Events

Aditya Chinchure, Sahithya Ravi, Raymond Ng +3

The commonsense reasoning capabilities of vision-language models (VLMs), especially in abductive reasoning and defeasible reasoning, remain poorly understood. Most benchmarks focus…

cs.CV2025

The Power of One: A Single Example is All it Takes for Segmentation in VLMs

Mir Rayat Imtiaz Hossain, Mennatullah Siam, Leonid Sigal +1

Large-scale vision-language models (VLMs), trained on extensive datasets of image-text pairs, exhibit strong multimodal understanding capabilities by implicitly learning associatio…