1 citations · 1 across the 8 of their papers we have counts for
4 papers · 1 filter
Cross-Cultural Value Attribution in Large Vision-Language Models
Phillip Howard, Xin Su, Kathleen C. Fraser
The rapid adoption of large vision-language models (LVLMs) in recent years has been accompanied by growing fairness concerns due to their propensity to reinforce harmful societal s…
Cultural Counterfactuals: Evaluating Cultural Biases in Large Vision-Language Models with Counterfactual Examples
Phillip Howard, Xin Su, Kathleen C. Fraser
Large Vision-Language Models (LVLMs) have grown increasingly powerful in recent years, but can also exhibit harmful biases. Prior studies investigating such biases have primarily f…
Geometry-Aware CLIP Retrieval via Local Cross-Modal Alignment and Steering
Nirmalendu Prakash, Narmeen Fatimah Oozeer, Xin Su +8
CLIP retrieval is typically framed as a pointwise similarity problem in a shared embedding space. While CLIP achieves strong global cross-modal alignment, many retrieval failures a…
Training-Free Mitigation of Language Reasoning Degradation After Multimodal Instruction Tuning
Neale Ratzlaff, Man Luo, Xin Su +2
Multimodal models typically combine a powerful large language model (LLM) with a vision encoder and are then trained on multimodal data via instruction tuning. While this process a…