7 papers · 1 filter
How and What to Imagine? Visual Thinking in Unified Multimodal Models for Cross-View Spatial Reasoning
Qian Yang, Ankur Sikarwar, Huy Le +4
Cross-view spatial reasoning remains a weak spot for vision-language models (VLMs): they often reason in language and lose the fine-grained geometry needed for the task. Thinking w…
DynFrame: Adaptive Reasoning-Driven Multimodal Framework with Dynamic Frame Augmentation for Complex Video Understanding
Peng Zhang, Guanghao Zhang, Wanggui He +10
Recent video multimodal large language models (MLLMs) increasingly couple step-by-step reasoning with on-demand visual evidence retrieval, allowing models to revisit relevant video…
RiT: Vanilla Diffusion Transformers Suffice in Representation Space
Le Zhang, Ning Mang, Aishwarya Agrawal
Flow matching with -prediction -- regressing the clean data point rather than the ambient velocity -- is known to exploit low-dimensional manifold structure effectively in pixel…
Assessing and Learning Alignment of Unimodal Vision and Language Models
Le Zhang, Qian Yang, Aishwarya Agrawal
How well are unimodal vision and language models aligned? Although prior work have approached answering this question, their assessment methods do not directly translate to how the…
Investigating Prompting Techniques for Zero- and Few-Shot Visual Question Answering
Rabiul Awal, Le Zhang, Aishwarya Agrawal
In this paper, we explore effective prompting techniques to enhance zero- and few-shot Visual Question Answering (VQA) performance in contemporary Vision-Language Models (VLMs). Ce…
VisMin: Visual Minimal-Change Understanding
Rabiul Awal, Saba Ahmadi, Le Zhang +1
Fine-grained understanding of objects, attributes, and relationships between objects is crucial for visual-language models (VLMs). Existing benchmarks primarily focus on evaluating…