11 citations · 12 across the 2 of their papers we have counts for
2 papers
cs.CV2024★ 1 cited
MuirBench: A Comprehensive Benchmark for Robust Multi-image Understanding
Fei Wang, Xingyu Fu, James Y. Huang +18
We introduce MuirBench, a comprehensive benchmark that focuses on robust multi-image understanding capabilities of multimodal LLMs. MuirBench consists of 12 diverse multi-image tas…
cs.CV2023★ 11 cited
Improving Zero-Shot Generalization for CLIP with Synthesized Prompts
Zhengbo Wang, Jian Liang, Ran He +3
With the growing interest in pretrained vision-language models like CLIP, recent research has focused on adapting these models to downstream tasks. Despite achieving promising resu…