2 papers
cs.CV2025
RSVLM-QA: A Benchmark Dataset for Remote Sensing Vision Language Model-based Question Answering
Xing Zi, Jinghao Xiao, Yunxiao Shi +4
Visual Question Answering (VQA) in remote sensing (RS) is pivotal for interpreting Earth observation data. However, existing RS VQA datasets are constrained by limitations in annot…
cs.CV2024
Tangram: Benchmark for Evaluating Geometric Element Recognition in Large Multimodal Models
Chao Zhang, Jiamin Tang, Jing Xiao
Significant advancements in Large Multimodal Models (LMMs) have enabled them to tackle complex problems involving visual-mathematical reasoning. However, their ability to identify…