2 citations · 2 across the 1 of their papers we have counts for
Showing cs.CVShow all
2 papers · 1 filter
cs.CV2024★ 2 cited
Open-ended VQA benchmarking of Vision-Language models by exploiting Classification datasets and their semantic hierarchy
Simon Ging, María A. Bravo, Thomas Brox
The evaluation of text-generative vision-language models is a challenging yet crucial endeavor. By addressing the limitations of existing Visual Question Answering (VQA) benchmarks…
cs.CV2019
MAIN: Multi-Attention Instance Network for Video Segmentation
Juan Leon Alcazar, Maria A. Bravo, Ali K. Thabet +4
Instance-level video segmentation requires a solid integration of spatial and temporal information. However, current methods rely mostly on domain-specific information (online lear…