1 citations · 1 across the 5 of their papers we have counts for
Showing cs.CVShow all
2 papers · 1 filter
cs.CV2026
When Vision Speaks for Sound
Xiaofei Wen, Wenjie Jacky Mo, Xingyu Fu +6
Despite rapid progress in video-capable MLLMs, we find that their apparent audio understanding in videos is often vision-driven: models rely on visual cues to infer or hallucinate…
cs.CV2024
MuirBench: A Comprehensive Benchmark for Robust Multi-image Understanding
Fei Wang, Xingyu Fu, James Y. Huang +18
We introduce MuirBench, a comprehensive benchmark that focuses on robust multi-image understanding capabilities of multimodal LLMs. MuirBench consists of 12 diverse multi-image tas…