4 papers
Doc-CoB: Enhancing Document Understanding with Visual Chain-of-Boxes Reasoning
Ye Mo, Kai Ye, Xianwei Mao +9
Document understanding aims to perform question answering and information extraction over document images, where the visual content is highly information-dense and most queries rel…
Can Multimodal Large Language Models Truly Understand Small Objects?
Fujun Han, Junan Chen, Xintong Zhu +4
Multimodal Large Language Models (MLLMs) have shown promising potential in diverse understanding tasks, e.g., image and video analysis, math and physics olympiads. However, they re…
AstroMMBench: A Benchmark for Evaluating Multimodal Large Language Models Capabilities in Astronomy
Jinghang Shi, Xiaoyu Tang, Yang Huang +4
Astronomical image interpretation presents a significant challenge for applying multimodal large language models (MLLMs) to specialized scientific tasks. Existing benchmarks focus…
StructChart: On the Schema, Metric, and Augmentation for Visual Chart Understanding
Renqiu Xia, Haoyang Peng, Hancheng Ye +7
Charts are common in literature across various scientific fields, conveying rich information easily accessible to readers. Current chart-related tasks focus on either chart percept…