Showing cs.CVShow all
2 papers · 1 filter
cs.CV2026
Unlocking ImageNet's Multi-Object Nature: Automated Large-Scale Multilabel Annotation
Junyu Chen, Md Yousuf Harun, Christopher Kanan
The original ImageNet benchmark enforces a single-label assumption, despite many images depicting multiple objects. This leads to label noise and limits the richness of the learnin…
cs.CV2024
BloomVQA: Assessing Hierarchical Multi-modal Comprehension
Yunye Gong, Robik Shrestha, Jared Claypoole +4
We propose a novel VQA dataset, BloomVQA, to facilitate comprehensive evaluation of large vision-language models on comprehension tasks. Unlike current benchmarks that often focus…