Showing cs.CVShow all
2 papers · 1 filter
cs.CV2026
Can Multimodal Large Language Models Truly Understand Small Objects?
Fujun Han, Junan Chen, Xintong Zhu +4
Multimodal Large Language Models (MLLMs) have shown promising potential in diverse understanding tasks, e.g., image and video analysis, math and physics olympiads. However, they re…
cs.CV2024
StructChart: On the Schema, Metric, and Augmentation for Visual Chart Understanding
Renqiu Xia, Haoyang Peng, Hancheng Ye +7
Charts are common in literature across various scientific fields, conveying rich information easily accessible to readers. Current chart-related tasks focus on either chart percept…