#benchmark dataset

29 results
cs.AI2026

PathView-Bench: Can Multimodal Large Language Models Achieve Fine-grained Multiscale Understanding of Pathology Images?

Zongyi Chen, Yu Liang, Jie Lin +1

The paper presents PathVU, a benchmark that tests multimodal large language models on fine-grained, multiscale visual understanding of pathology images using region- and slide-leve…

#multimodal large language models#pathology imaging#visual question answering#multiscale understanding
cs.CV2026

LoMeVQA: A Comprehensive Benchmark for Longitudinal Medical VQA

Zhilin Wu, Zhangkai Ni, Chengmei Yang +4

The paper introduces LoMeVQA, a large benchmark of 206K longitudinal medical visual question answering pairs designed to evaluate temporal reasoning over sequential medical images,…

#longitudinal analysis#medical visual question answering#temporal reasoning#multimodal language models
cs.CV2026

ARD-REFSM: Enhancing Reflection Symmetry Detection with Asymmetric Denoising and Rotation Equivariance

Dongfu Yin, Rourou Su, Cong Zhao +1

The paper introduces a method that removes asymmetric background clutter and enforces rotation-equivariant feature matching to improve detection of reflection symmetry in images, a…

#reflection symmetry detection#rotation equivariance#asymmetric denoising#deep learning
cs.CV2026

SciFigQual-Bench: A Benchmark for Scientific Figure Quality Assessment with Full-Manuscript Context

Zihan Deng, Chuanzhi Xu, Huiqi Liang +3

The paper introduces SciFigQual-Bench, a benchmark dataset that evaluates the quality of scientific figures within full manuscript context across five dimensions, and presents a cr…

#scientific figure quality assessment#cross-modal evaluation#benchmark dataset#contextual relevance
cs.CV2026

ICDAR 2026 Competition on Information Extraction from Atomic Layer Deposition/Etching (ALD/E) Scientific Figures

Fahad Ahmed, Sören Auer, Jennifer D'Souza

The paper presents the ICDAR 2026 competition and the Sci-ImageMiner benchmark for extracting and reasoning over information in scientific figures related to atomic layer depositio…

#scientific figure understanding#multimodal learning#information extraction#visual question answering
cs.CR2026

A Controlled Candidate-Set Benchmark for Offline Satellite-Security Plan Decomposition

João Paolo Cavalcante Martins Oliveira, Lucas Teske, Paulo Matias

The paper introduces a benchmark for offline satellite-security plan decomposition, providing a case‑disjoint dataset and a low‑rank decomposition adapter, and evaluates selection…

#satellite security#plan decomposition#benchmark dataset#low-rank adapter
cs.LG2026

PowerAtlas: Towards Electricity-Computing Co-Scheduling for Power Systems

Kaiwen Jiang, Siya Xu, Ziyue Zhu +3

PowerAtlas is an LLM‑agent framework that jointly schedules electricity supply and computing workloads in data centers, ensuring grid operational constraints and computing service…

#electricity-computing co-scheduling#llm agents#grid constraints#data center scheduling
cs.CR2026

HoF-Bench: Rediscovering Real AI-Discovered CVEs Without Frontier Models

Petr Simecek, Elnaz Babayeva, Jiri Balhar +23

The paper presents HoF-Bench, a benchmark of 95 AI‑discovered CVEs from open‑source projects, and evaluates several LLM‑based vulnerability detectors, finding that only a minimal a…

#vulnerability detection#large language models#benchmark dataset#software security
cs.AI2026

Fewer Clarifications, Better Code: Benchmarking Cross-Session Personalized Ambiguity Adaptation in Coding Assistants

Zijian Xu, Wenshuo Zhang, Zisen Qin +4

The paper defines personalized ambiguity adaptation for coding assistants, introduces the CAPA benchmark to evaluate how well models use a user's past resolved sessions to handle r…

#code generation#personalized assistance#ambiguity resolution#user history
cs.CV2026

PanDent: Toward Comprehensive Tooth-Level Structure-Language Consistency in Dental Radiology

Xiaohan Li, Xinyu Liu, Chang Liu +4

The paper presents PanDent, a large-scale benchmark of dental panoramic radiographs with expert-validated tooth-level annotations and corresponding radiology reports, designed to e…

#dental radiology#multimodal language models#tooth-level annotation#benchmark dataset
cs.CL2026

WorkSurface-Bench: Benchmarking Enterprise Agents on Multi-Surface Knowledge Routing

Hao Liang, Meiyi Qiang, Sizhe Qiu +2

The paper introduces WorkSurface-Bench, a benchmark that tests enterprise agents' ability to select the correct knowledge source (documents, tables, or graphs) before answering que…

#knowledge routing#enterprise agents#multi-surface retrieval#benchmark dataset
cs.CV2026

ObliCity: A Benchmark and Baseline for Roof-to-Ground Projection Displacement Correction

Kai Li, Yupeng Deng, Ligao Deng +6

The paper presents ObliCity, a large-scale benchmark for extracting roof-to-footprint offset vectors in oblique urban remote sensing images, and introduces DragRoof, an ODE-based m…

#oblique aerial imagery#roof-to-footprint alignment#geometric correction#benchmark dataset
cs.CV2026

BG-REAL: A Public Real-Data Anchored Benchmark for Background Manipulation Detection and Localization

Bugra Alperen Uluirmak, Rifat Kurban

The paper introduces BG-REAL, a publicly available benchmark for detecting and localizing manipulations that occur in the background of images, built from real Open Images data and…

#image forensics#background manipulation#benchmark dataset#detection
cs.AI2026

MedRealMM: A Real-World Multimodal Benchmark for Chinese Online Medical Consultation

Runhan Shi, Quan Zhou, Yuqian Xu +14

The paper presents MedRealMM, a large-scale benchmark of real Chinese online medical consultations that includes both text and patient-uploaded images, and evaluates how well large…

#multimodal learning#medical consultation#large language models#clinical evaluation
cs.CL2026

SciDiagramEdit: Learning to Edit Scientific Diagrams from Paper Revisions

Yasheng Sun, Zezi Zeng, Yifan Yang +4

The paper introduces SciDiagramEdit, a system that learns to edit scientific figures based on natural-language revision instructions by training on before-and-after figure pairs fr…

#figure editing#instruction following#vector graphics#skill evolution
cs.CL2026

Live Gurbani Tracking: A Benchmark and Reference System for Captioning Sikh Kirtan

Karanbir Singh

The paper introduces a benchmark and a reference system for live, closed‑vocabulary captioning of Sikh Kirtan, providing annotated recordings, a scoring metric, and an on‑device mo…

#kirtan captioning#closed-vocabulary transcription#live speech recognition#gurmukhi script
cs.CV2026

Towards Spatial Supersensing in the Wild

Tianjun Gu, Tianyu Xin, Kuan Zhang +12

The paper introduces VSI‑Super‑Wild, a large benchmark of real‑world long videos with human‑verified QA pairs to evaluate how well multimodal models can track and reason about agen…

#spatial reasoning#long-term video understanding#multimodal world modeling#benchmark dataset
cs.MA2026

An Explainable Agentic System for Detection of Conversational Scams with Summary-Based Memory

Ahmed Omar Salim Adnan, Yogananda Manjunath, Shivanjali Khare

The paper presents an explainable, agentic system that uses summary‑based memory to detect multi‑turn conversational scams, introduces a new benchmark dataset (ConScamBench‑278), a…

#conversational scam detection#explainable AI#agentic systems#benchmark dataset
cs.CV2026

OccTrack360: 4D Panoptic Occupancy Tracking from Surround-View Fisheye Cameras

Yongzhi Lin, Kai Luo, Yuanfan Zheng +4

The paper introduces OccTrack360, a benchmark for 4D panoptic occupancy tracking using surround-view fisheye cameras, and proposes FoSOcc, a baseline framework with modules to hand…

#panoptic occupancy tracking#fisheye cameras#surround-view perception#3d scene understanding
cs.CL2026

UXBench: Benchmarking User Experience in AI Assistants

Mengze Hong, Xia Zeng, Zeyang Lei +26

UXBench is a user‑centric benchmark that uses real interaction logs to evaluate how well AI assistants align with user preferences and generate engaging dialogue, featuring three t…

#user experience evaluation#ai assistants#dialogue systems#benchmark dataset
cs.CV2026

MoHallBench: A Benchmark for Motion Hallucination in Video Large Language Models

Sihan Chen, Jiale Li, Jianghang Lin +1

The paper introduces MoHallBench, a large benchmark designed to evaluate and diagnose motion hallucination—incorrectly inferred human motions—in video large language models, coveri…

#motion hallucination#video large language models#benchmark dataset#video understanding
cs.SE2026

Deep4ge: DNN Training Trajectories for Fault Detection and Diagnosis

Sigma Jahan

The paper introduces Deep4ge, a benchmark dataset of over 14,000 deep neural network training runs—including faulty and correct variants—capturing per‑epoch metrics and features to…

#fault detection#deep learning#training trajectories#benchmark dataset
cs.CL2026

HiQA: A Hierarchical Contextual Augmentation RAG for Multi-Documents QA

Xinyue Chen, Pengyu Gao, Jiangjiang Song +1

HiQA is a framework for multi‑document question answering that uses hierarchical contextual augmentation and a multi‑route retrieval mechanism with cascading metadata to improve re…

#multi-document question answering#retrieval-augmented generation#hierarchical augmentation#metadata integration
cs.CV2026

Towards Temporal Compositional Reasoning in Long-Form Sports Videos

Siyu Cao, Lu Zhang, Ruizhe Zeng +1

The paper introduces SportsTime, a large benchmark of long-form sports videos with detailed temporal evidence annotations, and proposes the Chain-of-Time Reasoning (CoTR) framework…

#sports video analysis#temporal compositional reasoning#multimodal large language models#benchmark dataset
← Prev1 / 2Next →