computer vision

SciFigQual-Bench: A Benchmark for Scientific Figure Quality Assessment with Full-Manuscript Context

arXiv:2607.27084

summary

The paper introduces SciFigQual-Bench, a benchmark dataset that evaluates the quality of scientific figures within full manuscript context across five dimensions, and presents a cross‑modal evaluation system (SFQ‑Agent) that leverages large language models to score these figures.

Abstract

Scientific images are the core elements of presenting experimental conclusions, elaborating system architecture, and supporting comparative arguments in scientific papers. However, existing image quality assessment (IQA) methods are predominantly designed for natural photographs or AI-generated content, which cannot be directly applied to scientific papers. The few existing studies on scholarly charts remain confined to visual-surface comparisons, failing to verify caption alignment, citation relevance, or visual misleadingness. To address this, we propose SciFigQual-Bench, a full-text contextual benchmark that evaluates scientific images across five dimensions (clarity, layout, caption fit, context relevance, and misleading risk). The data covers top computer-science conferences from 2020 to 2025; 6,308 images were independently scored by multiple domain experts in five dimensions and aggregated into gold-standard annotations. Unlike previous scientific figure benchmarks, our dataset binds each image to its caption, citing sentence, and manuscript context. To enable automated evaluation on this benchmark, we designed a staged cross-modal evaluation framework SFQ-Agent to achieve auditable and refined scoring through the collection and fusion of modal evidence. Multiple mainstream large models were evaluated on the test subset eval1200, and SFQ-Agent (F3) equipped with GPT-5.6-Sol achieved the lowest overall average absolute error (0.418) and the highest consistency rate (93.4%), consistently outperforming both direct evaluation and auxiliary (Sidecar) visual language model evaluation schemes.

†Equal contribution. Affiliations: 1: The University of Hong Kong 2: The University of Sydney 3: University of Electronic Science and Technology of China Corresponding authors: Zihan Deng ([email protected]), Chuanzhi Xu ([email protected]) Project page: https://frankdengai.github.io/SciFigQual-Bench Source code & dataset: https://github.com/FrankDengAI/SciFigQual-Bench

Topics & keywords

#scientific figure quality assessment#cross-modal evaluation#benchmark dataset#contextual relevance#large language modelsSciFigQual-BenchSFQ-AgentGPT-5.6-Solfigure claritycaption fitmisleading riskmodal evidence fusion