3 papers
stat.ML2025
Multimodal Datasets with Controllable Mutual Information
Raheem Karim Hashmani, Garrett W. Merz, Helen Qu +2
We introduce a framework for generating highly multimodal datasets with explicitly calculable mutual information (MI) between modalities. This enables the construction of benchmark…
cs.CV2025
Impact of Pretraining Word Co-occurrence on Compositional Generalization in Multimodal Models
Helen Qu, Sang Michael Xie
CLIP and large multimodal models (LMMs) have better accuracy on examples involving concepts that are highly represented in the training data. However, the role of concept combinati…
cs.CV2024
Connect Later: Improving Fine-tuning for Robustness with Targeted Augmentations
Helen Qu, Sang Michael Xie
Models trained on a labeled source domain (e.g., labeled images from wildlife camera traps) often generalize poorly when deployed on an out-of-distribution (OOD) target domain (e.g…