8 papers
Generative Action Tell-Tales: Assessing Human Motion in Synthesized Videos
Xavier Thomas, Youngsun Lim, Ananya Srinivasan +2
Despite rapid advances in video generative models, robust metrics for evaluating visual and temporal correctness of complex human actions remain elusive. Critically, existing pure-…
Some Modalities are More Equal Than Others: Decoding and Architecting Multimodal Integration in MLLMs
Tianle Chen, Chaitanya Chakka, Arjun Reddy Akula +2
Despite remarkable advancements in Multimodal Large Language Models (MLLMs), a fundamental question remains: are MLLMs robust to contradicting modalities? To rigorously study this,…
GeoDE: a Geographically Diverse Evaluation Dataset for Object Recognition
Vikram V. Ramaswamy, Sing Yu Lin, Dora Zhao +4
Current dataset collection methods typically scrape large amounts of data from the web. While this technique is extremely scalable, data collected in this way tends to reinforce st…
: Interpreting and leveraging semantic information in diffusion models
Dahye Kim, Xavier Thomas, Deepti Ghadiyaram
We study rich visual semantic information is represented within various layers and denoising timesteps of different diffusion architectures. We uncover monosemantic…
Progressive Prompt Detailing for Improved Alignment in Text-to-Image Generative Models
Ketan Suhaas Saichandran, Xavier Thomas, Prakhar Kaushik +1
Text-to-image generative models often struggle with long prompts detailing complex scenes, diverse objects with distinct visual characteristics and spatial relationships. In this w…
Improving Physical Object State Representation in Text-to-Image Generative Systems
Tianle Chen, Chaitanya Chakka, Deepti Ghadiyaram
Current text-to-image generative models struggle to accurately represent object states (e.g., "a table without a bottle," "an empty tumbler"). In this work, we first design a fully…