17 papers
TECCI: Tricky Edits of Collected and Curated Images
Aishwarya Agrawal, Roy Hirsch, Yasumasa Onoe +2
Despite tremendous recent progress, current text-guided image editing methods still struggle with many aspects of editing involving instruction following, minimally editing the sou…
How and What to Imagine? Visual Thinking in Unified Multimodal Models for Cross-View Spatial Reasoning
Qian Yang, Ankur Sikarwar, Huy Le +4
Cross-view spatial reasoning remains a weak spot for vision-language models (VLMs): they often reason in language and lose the fine-grained geometry needed for the task. Thinking w…
RiT: Vanilla Diffusion Transformers Suffice in Representation Space
Le Zhang, Ning Mang, Aishwarya Agrawal
Flow matching with -prediction -- regressing the clean data point rather than the ambient velocity -- is known to exploit low-dimensional manifold structure effectively in pixel…
From Where Things Are to What They Are For: Benchmarking Spatial-Functional Intelligence in Multimodal LLMs
Le Zhang, Jihan Yang, Soundarya Krishnan +11
Human-level agentic intelligence extends beyond low-level geometric perception, evolving from recognizing where things are to understanding what they are for. While existing benchm…
Discovering Failure Modes in Vision-Language Models using RL
Kanishk Jain, Qian Yang, Shravan Nayak +3
Vision-language Models (VLMs), despite achieving strong performance on multimodal benchmarks, often misinterpret straightforward visual concepts that humans identify effortlessly,…
Communicating about Space: Language-Mediated Spatial Integration Across Partial Views
Ankur Sikarwar, Debangan Mishra, Sudarshan Nikhil +2
Humans build shared spatial understanding by communicating partial, viewpoint-dependent observations. We ask whether Multimodal Large Language Models (MLLMs) can do the same, align…