40 citations · 43 across the 6 of their papers we have counts for
Showing cs.CVShow all
3 papers · 1 filter
cs.CV2025
VER-Bench: Evaluating MLLMs on Reasoning with Fine-Grained Visual Evidence
Chenhui Qiang, Zhaoyang Wei, Xumeng Han +5
With the rapid development of MLLMs, evaluating their visual capabilities has become increasingly crucial. Current benchmarks primarily fall into two main types: basic perception b…
cs.CV2023★ 2 cited
ChartReader: A Unified Framework for Chart Derendering and Comprehension without Heuristic Rules
Zhi-Qi Cheng, Qi Dai, Siyao Li +3
Charts are a powerful tool for visually conveying complex data, but their comprehension poses a challenge due to the diverse chart types and intricate components. Existing chart co…
cs.CV2020★ 40 cited
Detecting Hate Speech in Multi-modal Memes
Abhishek Das, Japsimar Singh Wahi, Siyao Li
In the past few years, there has been a surge of interest in multi-modal problems, from image captioning to visual question answering and beyond. In this paper, we focus on hate sp…