#model evaluation

try —

19 papers match

physics.ao-ph2026

Weather Emulators at the Frontier of Heat Extremes Predictability

Cas Decancq, Thomas Mortier, Jessica Keune +1

The paper compares six modern deep‑learning weather emulators with traditional dynamical and statistical models for forecasting global near‑surface temperature and extreme heat at…

#weather forecasting#deep learning#heat extremes#extended-range prediction
cs.CL2026

Diagnosing Fine-Grained Inconsistency Classification in Financial Disclosure Text

Aman Kumar, Lasitha Vidyaratne, Dipanjan D Ghosh +2

The paper investigates fine-grained classification of different inconsistency types in financial disclosure texts, evaluating various supervised and large language models on a synt…

#financial disclosures#inconsistency detection#fine-grained classification#evidence extraction
cs.LG2026

BayesAME: Bayesian Active Model Evaluation

Paula Cordero Encinar, Taylan Cemgil, Arnaud Doucet +2

BayesAME is a Bayesian sequential framework that automatically determines the size of a coreset for evaluating large generative models, using latent ability models and information‑…

#model evaluation#active learning#bayesian inference#coreset selection
cs.CV2026

Object Detection for Autonomous Driving in Chinese Rural Scenes: An Experimental Study on Real-Synthetic Data Mixing and Model Evaluation

Danning Zhu, Ziyan Lin, Jing Wu

The paper introduces a mixed real‑synthetic dataset for object detection on Chinese rural roads and evaluates 13 popular detectors, finding that a moderate amount of synthetic data…

#autonomous driving#object detection#synthetic data#rural scenes
stat.ML2026

Can Deep Generative Models Reproduce Non-Stationary Gaussian Random Fields?

Daniel Kua, Yan Song

The paper evaluates several deep generative models on a synthetic non‑stationary Gaussian random field to see how well they recover the true mean and covariance, and demonstrates t…

#deep generative models#non-stationary gaussian random fields#spatial statistics#model evaluation
cs.AI2026

Benchmarking Multimodal Large Language Models for Scientific Visualization Literacy

Patrick Phuoc Do, Chau M. Ta, Chaoli Wang

The paper evaluates six multimodal large language models on a standardized scientific visualization literacy test, comparing their performance to human participants and highlightin…

#multimodal language models#scientific visualization#literacy benchmarking#model evaluation
cs.RO2026

JoyAI-Sim: A Simulation-Enabled Interconversion Toolchain for the Embodied Data Pyramid

Peidong Liu, Yongce Liu, Songyan Guo +34

JoyAI-Sim is a toolchain that connects real robots, simulation, and human demonstrations to enable scalable evaluation and generation of robot training data using calibrated digita…

#simulation#digital twins#human-robot interaction#data generation
cs.LG2026

Towards a Unified Multidimensional Explainability Metric: Evaluating Trustworthiness in AI Models

Georgios Makridis, Georgios Fatouros, Athanasios Kiourtis +4

The paper proposes a framework that evaluates explainability methods like LIME and SHAP across models and datasets using fidelity, simplicity, and stability, and builds a knowledge…

#explainable ai#model evaluation#trustworthiness#benchmarking
physics.ed-ph2026

Assessing AI in Introductory Physics Problem Solving

Amir Bralin, N. Sanjay Rebello

The paper evaluates OpenAI's o4-mini large language model on introductory physics problems from Halliday and Resnick, finding about 90% overall accuracy but lower performance on im…

#large language models#physics problem solving#introductory physics#image-text integration
cs.AI2026

The Model Knows Your Project, Not You: Measuring Recognition in LLMs with NameRank

Bojie Li, Noah Shi

The paper introduces NameRank, a metric that quantifies how well large language models can recognize specific people or tools from their internal weights without external retrieval…

#large language models#entity recognition#knowledge probing#model evaluation
cs.CR2026

VanillaBench: The Hidden Accuracy Cost of Adversarial Robustness

Niklas Bunzel

The paper presents VanillaBench, a benchmark that measures how much clean (vanilla) accuracy is lost when models are trained for adversarial robustness, revealing a larger accuracy…

#adversarial robustness#benchmarking#clean accuracy gap#model evaluation
cs.CL2026

Epistemic Stance Flexibility Probing: Measuring Prompt-Conditioned Register Shift in Large Language Models

Binwen Liu, Yilin Ren

The paper introduces ESFP, a benchmark that tests whether large language models can shift between neutral attribution and personal stance when prompted differently, and evaluates t…

#epistemic stance#prompt conditioning#large language models#benchmarking
cs.CV2026

SpurLens: Automatic Detection of Spurious Cues in Multimodal LLMs

Parsa Hosseini, Sumit Nawathe, Mazda Moayeri +2

The paper introduces SpurLens, an automated pipeline that uses GPT-4 and open-set object detectors to find spurious visual cues in multimodal large language models, showing that th…

#multimodal language models#spurious correlations#visual bias#prompt engineering
cs.LG2026

Detecting Explanatory Insufficiency in Learned Representations: A Framework for Representational Vigilance

Jacques Raynal, Pierre Slangen, Elsa Raynal +1

The paper proposes VER, a conceptual framework for monitoring learned representations to detect unexplained residual structures that standard performance metrics miss, offering a d…

#representation learning#model evaluation#explanatory insufficiency#diagnostic frameworks
cs.LG2026

Trustworthy synthetic data for campaign decision support: strategy simulation fidelity and the PolicySynth framework

Tung Dang, The Hung Phung, Son Lam Nguyen +1

The paper presents PolicySynth, a synthetic data generation framework designed to preserve decision alignment for marketing campaign support, and introduces strategy simulation fid…

#synthetic data#decision support systems#campaign optimization#privacy
cs.CL2026

Characterising AI Models for Cataloguing

Miguel Arana-Catania, Neil Jefferies

The paper evaluates various AI models for automatically generating catalogue records for digitized collections, comparing implementations through qualitative and quantitative exper…

#cataloguing#digital collections#metadata generation#model evaluation
cs.AI2026

Where Experts Disagree, Models Fail: Detecting Implicit Legal Citations in French Court Decisions

Avrile Floro, Tamara Dhorasoo, Soline Pellez +1

The paper introduces a benchmark for detecting implicit citations of the French Civil Code in court decisions and shows that cases where legal experts disagree are especially hard…

#legal citation detection#implicit reasoning#expert disagreement#benchmark dataset
cs.CV2026

Are DeepFakes Realistic Enough? Exploring Semantic Mismatch as a Novel Challenge

Sharayu Nilesh Deshmukh, Kailash A. Hambarde, Joana C. Costa +2

The paper introduces a new evaluation class for DeepFake detection that captures semantic inconsistencies between audio and video (RARV‑SMM) and tests how current models handle suc…

#deepfake detection#semantic mismatch#audio‑visual forensics#multimodal robustness
cs.CL2026

Evaluating Nonuniform Dependability Across Response Conditions: A Conditional Generalizability Framework Illustrated in Automated Essay Scoring

Yi Gui

The paper proposes a conditional generalizability framework to assess how reliable automated essay scoring systems are across different response conditions, using entropy-based str…

#automated essay scoring#generalizability theory#dependability analysis#entropy stratification

One search, two signals: results blend meaning (embedding similarity, so papers that never use your words still surface) with keyword matches on titles, abstracts and summaries. Free, no sign-in needed.