5 papers
Semantic Richness or Geometric Reasoning? The Fragility of VLM's Visual Invariance
Jason Qiu, Zachary Meurer, Xavier Thomas +1
This work investigates the fundamental fragility of state-of-the-art Vision-Language Models (VLMs) under basic geometric transformations. While modern VLMs excel at semantic tasks…
Generative Action Tell-Tales: Assessing Human Motion in Synthesized Videos
Xavier Thomas, Youngsun Lim, Ananya Srinivasan +2
Despite rapid advances in video generative models, robust metrics for evaluating visual and temporal correctness of complex human actions remain elusive. Critically, existing pure-…
Some Modalities are More Equal Than Others: Decoding and Architecting Multimodal Integration in MLLMs
Tianle Chen, Chaitanya Chakka, Arjun Reddy Akula +2
Despite remarkable advancements in Multimodal Large Language Models (MLLMs), a fundamental question remains: are MLLMs robust to contradicting modalities? To rigorously study this,…
Progressive Prompt Detailing for Improved Alignment in Text-to-Image Generative Models
Ketan Suhaas Saichandran, Xavier Thomas, Prakhar Kaushik +1
Text-to-image generative models often struggle with long prompts detailing complex scenes, diverse objects with distinct visual characteristics and spatial relationships. In this w…
What's in a Latent? Leveraging Diffusion Latent Space for Domain Generalization
Xavier Thomas, Deepti Ghadiyaram
Domain Generalization aims to develop models that can generalize to novel and unseen data distributions. In this work, we study how model architectures and pre-training objectives…