works on

From the 1 of 16 linked papers with an AI index.

activity
20242026
most citedScaling Knowledge Graph Construction through Synthetic Data Generation and Distillation

1 citations · 1 across the 10 of their papers we have counts for

collaborators
Showing cs.CVShow all

5 papers · 1 filter

cs.CV2026

Improving Sample Diversity in Autoregressive Text-to-Image Generation via Cluster Truncation

Trang Nguyen, Shuang Wu, Runyan Tan +1

While diffusion models achieve state-of-the-art image quality for text-to-image (T2I) generation, recent work has demonstrated that they suffer from sample diversity collapse. In t…

cs.CV2026

Cross-Cultural Value Attribution in Large Vision-Language Models

Phillip Howard, Xin Su, Kathleen C. Fraser

The rapid adoption of large vision-language models (LVLMs) in recent years has been accompanied by growing fairness concerns due to their propensity to reinforce harmful societal s…

cs.CV2026

Cultural Counterfactuals: Evaluating Cultural Biases in Large Vision-Language Models with Counterfactual Examples

Phillip Howard, Xin Su, Kathleen C. Fraser

Large Vision-Language Models (LVLMs) have grown increasingly powerful in recent years, but can also exhibit harmful biases. Prior studies investigating such biases have primarily f…

cs.CV2026

Geometry-Aware CLIP Retrieval via Local Cross-Modal Alignment and Steering

Nirmalendu Prakash, Narmeen Fatimah Oozeer, Xin Su +8

CLIP retrieval is typically framed as a pointwise similarity problem in a shared embedding space. While CLIP achieves strong global cross-modal alignment, many retrieval failures a…

cs.CV2024

Training-Free Mitigation of Language Reasoning Degradation After Multimodal Instruction Tuning

Neale Ratzlaff, Man Luo, Xin Su +2

Multimodal models typically combine a powerful large language model (LLM) with a vision encoder and are then trained on multimodal data via instruction tuning. While this process a…