#benchmark dataset
29 resultsPathView-Bench: Can Multimodal Large Language Models Achieve Fine-grained Multiscale Understanding of Pathology Images?
Zongyi Chen, Yu Liang, Jie Lin +1
The paper presents PathVU, a benchmark that tests multimodal large language models on fine-grained, multiscale visual understanding of pathology images using region- and slide-leve…
LoMeVQA: A Comprehensive Benchmark for Longitudinal Medical VQA
Zhilin Wu, Zhangkai Ni, Chengmei Yang +4
The paper introduces LoMeVQA, a large benchmark of 206K longitudinal medical visual question answering pairs designed to evaluate temporal reasoning over sequential medical images,…
ARD-REFSM: Enhancing Reflection Symmetry Detection with Asymmetric Denoising and Rotation Equivariance
Dongfu Yin, Rourou Su, Cong Zhao +1
The paper introduces a method that removes asymmetric background clutter and enforces rotation-equivariant feature matching to improve detection of reflection symmetry in images, a…
SciFigQual-Bench: A Benchmark for Scientific Figure Quality Assessment with Full-Manuscript Context
Zihan Deng, Chuanzhi Xu, Huiqi Liang +3
The paper introduces SciFigQual-Bench, a benchmark dataset that evaluates the quality of scientific figures within full manuscript context across five dimensions, and presents a cr…
ICDAR 2026 Competition on Information Extraction from Atomic Layer Deposition/Etching (ALD/E) Scientific Figures
Fahad Ahmed, Sören Auer, Jennifer D'Souza
The paper presents the ICDAR 2026 competition and the Sci-ImageMiner benchmark for extracting and reasoning over information in scientific figures related to atomic layer depositio…
A Controlled Candidate-Set Benchmark for Offline Satellite-Security Plan Decomposition
João Paolo Cavalcante Martins Oliveira, Lucas Teske, Paulo Matias
The paper introduces a benchmark for offline satellite-security plan decomposition, providing a case‑disjoint dataset and a low‑rank decomposition adapter, and evaluates selection…
PowerAtlas: Towards Electricity-Computing Co-Scheduling for Power Systems
Kaiwen Jiang, Siya Xu, Ziyue Zhu +3
PowerAtlas is an LLM‑agent framework that jointly schedules electricity supply and computing workloads in data centers, ensuring grid operational constraints and computing service…
HoF-Bench: Rediscovering Real AI-Discovered CVEs Without Frontier Models
Petr Simecek, Elnaz Babayeva, Jiri Balhar +23
The paper presents HoF-Bench, a benchmark of 95 AI‑discovered CVEs from open‑source projects, and evaluates several LLM‑based vulnerability detectors, finding that only a minimal a…
Fewer Clarifications, Better Code: Benchmarking Cross-Session Personalized Ambiguity Adaptation in Coding Assistants
Zijian Xu, Wenshuo Zhang, Zisen Qin +4
The paper defines personalized ambiguity adaptation for coding assistants, introduces the CAPA benchmark to evaluate how well models use a user's past resolved sessions to handle r…
PanDent: Toward Comprehensive Tooth-Level Structure-Language Consistency in Dental Radiology
Xiaohan Li, Xinyu Liu, Chang Liu +4
The paper presents PanDent, a large-scale benchmark of dental panoramic radiographs with expert-validated tooth-level annotations and corresponding radiology reports, designed to e…
WorkSurface-Bench: Benchmarking Enterprise Agents on Multi-Surface Knowledge Routing
Hao Liang, Meiyi Qiang, Sizhe Qiu +2
The paper introduces WorkSurface-Bench, a benchmark that tests enterprise agents' ability to select the correct knowledge source (documents, tables, or graphs) before answering que…
ObliCity: A Benchmark and Baseline for Roof-to-Ground Projection Displacement Correction
Kai Li, Yupeng Deng, Ligao Deng +6
The paper presents ObliCity, a large-scale benchmark for extracting roof-to-footprint offset vectors in oblique urban remote sensing images, and introduces DragRoof, an ODE-based m…
BG-REAL: A Public Real-Data Anchored Benchmark for Background Manipulation Detection and Localization
Bugra Alperen Uluirmak, Rifat Kurban
The paper introduces BG-REAL, a publicly available benchmark for detecting and localizing manipulations that occur in the background of images, built from real Open Images data and…
MedRealMM: A Real-World Multimodal Benchmark for Chinese Online Medical Consultation
Runhan Shi, Quan Zhou, Yuqian Xu +14
The paper presents MedRealMM, a large-scale benchmark of real Chinese online medical consultations that includes both text and patient-uploaded images, and evaluates how well large…
SciDiagramEdit: Learning to Edit Scientific Diagrams from Paper Revisions
Yasheng Sun, Zezi Zeng, Yifan Yang +4
The paper introduces SciDiagramEdit, a system that learns to edit scientific figures based on natural-language revision instructions by training on before-and-after figure pairs fr…
Live Gurbani Tracking: A Benchmark and Reference System for Captioning Sikh Kirtan
Karanbir Singh
The paper introduces a benchmark and a reference system for live, closed‑vocabulary captioning of Sikh Kirtan, providing annotated recordings, a scoring metric, and an on‑device mo…
Towards Spatial Supersensing in the Wild
Tianjun Gu, Tianyu Xin, Kuan Zhang +12
The paper introduces VSI‑Super‑Wild, a large benchmark of real‑world long videos with human‑verified QA pairs to evaluate how well multimodal models can track and reason about agen…
An Explainable Agentic System for Detection of Conversational Scams with Summary-Based Memory
Ahmed Omar Salim Adnan, Yogananda Manjunath, Shivanjali Khare
The paper presents an explainable, agentic system that uses summary‑based memory to detect multi‑turn conversational scams, introduces a new benchmark dataset (ConScamBench‑278), a…
OccTrack360: 4D Panoptic Occupancy Tracking from Surround-View Fisheye Cameras
Yongzhi Lin, Kai Luo, Yuanfan Zheng +4
The paper introduces OccTrack360, a benchmark for 4D panoptic occupancy tracking using surround-view fisheye cameras, and proposes FoSOcc, a baseline framework with modules to hand…
UXBench: Benchmarking User Experience in AI Assistants
Mengze Hong, Xia Zeng, Zeyang Lei +26
UXBench is a user‑centric benchmark that uses real interaction logs to evaluate how well AI assistants align with user preferences and generate engaging dialogue, featuring three t…
MoHallBench: A Benchmark for Motion Hallucination in Video Large Language Models
Sihan Chen, Jiale Li, Jianghang Lin +1
The paper introduces MoHallBench, a large benchmark designed to evaluate and diagnose motion hallucination—incorrectly inferred human motions—in video large language models, coveri…
Deep4ge: DNN Training Trajectories for Fault Detection and Diagnosis
Sigma Jahan
The paper introduces Deep4ge, a benchmark dataset of over 14,000 deep neural network training runs—including faulty and correct variants—capturing per‑epoch metrics and features to…
HiQA: A Hierarchical Contextual Augmentation RAG for Multi-Documents QA
Xinyue Chen, Pengyu Gao, Jiangjiang Song +1
HiQA is a framework for multi‑document question answering that uses hierarchical contextual augmentation and a multi‑route retrieval mechanism with cascading metadata to improve re…
Towards Temporal Compositional Reasoning in Long-Form Sports Videos
Siyu Cao, Lu Zhang, Ruizhe Zeng +1
The paper introduces SportsTime, a large benchmark of long-form sports videos with detailed temporal evidence annotations, and proposes the Chain-of-Time Reasoning (CoTR) framework…