Showing cs.CVShow all
2 papers · 1 filter
cs.CV2025
Generalizing from SIMPLE to HARD Visual Reasoning: Can We Mitigate Modality Imbalance in VLMs?
Simon Park, Abhishek Panigrahi, Yun Cheng +3
Vision Language Models (VLMs) are impressive at visual question answering and image captioning. But they underperform on multi-step visual reasoning -- even compared to LLMs on the…
cs.CV2024
ConceptMix: A Compositional Image Generation Benchmark with Controllable Difficulty
Xindi Wu, Dingli Yu, Yangsibo Huang +2
Compositionality is a critical capability in Text-to-Image (T2I) models, as it reflects their ability to understand and combine multiple concepts from text descriptions. Existing e…