13 papers
On the Limitations of Cross-Lingual Consistency in Multilingual Text-to-image Generation
Sicheng Zhang, Zhonghao Yan, Binzhu Xie +4
Text-to-image (T2I) generation has achieved remarkable progress in recent years. However, existing research has largely focused on English-only settings, leaving cross-lingual perf…
ReasonCLIP-58M: Visually Grounded Commonsense Reasoning Supervision for CLIP
Sicheng Zhang, Muzammal Naseer, Binzhu Xie +5
CLIP and its variants are widely adopted visual backbones in multimodal systems, but their pretraining remains dominated by descriptive image-text alignment. As downstream applicat…
HumaniBench: A Human-Centric Framework for Large Multimodal Models Evaluation
Shaina Raza, Aravind Narayanan, Vahid Reza Khazaie +6
Although recent large multimodal models (LMMs) show impressive progress on vision language tasks, their alignment with human centered (HC) principles such as fairness, ethics, incl…
OncoTriad-QA: A Patient-Level Radiology-Pathology-Genomics Benchmark for Pan-Cancer Reasoning
Ahnaf Munir, Dannong Wang, Michael W. McDonald +3
Cancer diagnosis and characterization require integrating complementary evidence from radiology, pathology, genomics, and clinical metadata. However, most medical large language mo…
TIGeR: A Unified Framework for Time, Images and Geo-location Retrieval
David G. Shatwell, Sirnam Swetha, Mubarak Shah
Many real-world applications in digital forensics, urban monitoring, and environmental analysis require jointly reasoning about visual appearance, location, and time. Beyond standa…
BBQ-V: Benchmarking Visual Stereotype Bias in Large Multimodal Models
Vishal Narnaware, Ashmal Vayani, Rohit Gupta +2
Stereotype biases in Large Multimodal Models (LMMs) perpetuate harmful societal prejudices, undermining the fairness and equity of AI applications. As LMMs grow increasingly influe…