9 papers
Seeing Isn't Orienting: A Cognitively Informed Hierarchical Benchmark for Object Orientation in MLLMs
Nazia Tasnim, Keanu Nichols, Yuting Yan +4
Humans develop object orientation understanding progressively, from recognizing which way an object faces to reasoning about orientations across multiple objects. Yet existing visi…
Semantic Richness or Geometric Reasoning? The Fragility of VLM's Visual Invariance
Jason Qiu, Zachary Meurer, Xavier Thomas +1
This work investigates the fundamental fragility of state-of-the-art Vision-Language Models (VLMs) under basic geometric transformations. While modern VLMs excel at semantic tasks…
Swift Sampling: Selecting Temporal Surprises via Taylor Series
Dahye Kim, Bhuvan Sachdeva, Karan Uppal +3
While most frames in long-form video are redundant, the critical information resides in temporal surprises: moments where the actual visual features deviate from their predicted ev…
FAGER: Factually Grounded Evaluation and Refinement of Text-to-Image Models
Youngsun Lim, Cusuh Ham, Pin-Yu Chen +1
Existing text-to-image (T2I) evaluation metrics mainly assess whether generated images align with information explicitly stated in the prompt, but often fail to capture factual req…
FuTCR: Future-Targeted Contrast and Repulsion for Continual Panoptic Segmentation
Nicholas Ikechukwu, Keanu Nichols, Deepti Ghadiyaram +1
Continual Panoptic Segmentation (CPS) requires methods that can quickly adapt to new categories over time. The nature of this dense prediction task means that training images may c…
A Systematic Study of Cross-Modal Typographic Attacks on Audio-Visual Reasoning
Tianle Chen, Deepti Ghadiyaram
As audio-visual multi-modal large language models (MLLMs) are increasingly deployed in safety-critical applications, understanding their vulnerabilities is crucial. To this end, we…