#model evaluation
19 papers match
Weather Emulators at the Frontier of Heat Extremes Predictability
Cas Decancq, Thomas Mortier, Jessica Keune +1
The paper compares six modern deep‑learning weather emulators with traditional dynamical and statistical models for forecasting global near‑surface temperature and extreme heat at…
Diagnosing Fine-Grained Inconsistency Classification in Financial Disclosure Text
Aman Kumar, Lasitha Vidyaratne, Dipanjan D Ghosh +2
The paper investigates fine-grained classification of different inconsistency types in financial disclosure texts, evaluating various supervised and large language models on a synt…
BayesAME: Bayesian Active Model Evaluation
Paula Cordero Encinar, Taylan Cemgil, Arnaud Doucet +2
BayesAME is a Bayesian sequential framework that automatically determines the size of a coreset for evaluating large generative models, using latent ability models and information‑…
Object Detection for Autonomous Driving in Chinese Rural Scenes: An Experimental Study on Real-Synthetic Data Mixing and Model Evaluation
Danning Zhu, Ziyan Lin, Jing Wu
The paper introduces a mixed real‑synthetic dataset for object detection on Chinese rural roads and evaluates 13 popular detectors, finding that a moderate amount of synthetic data…
Can Deep Generative Models Reproduce Non-Stationary Gaussian Random Fields?
Daniel Kua, Yan Song
The paper evaluates several deep generative models on a synthetic non‑stationary Gaussian random field to see how well they recover the true mean and covariance, and demonstrates t…
Benchmarking Multimodal Large Language Models for Scientific Visualization Literacy
Patrick Phuoc Do, Chau M. Ta, Chaoli Wang
The paper evaluates six multimodal large language models on a standardized scientific visualization literacy test, comparing their performance to human participants and highlightin…
JoyAI-Sim: A Simulation-Enabled Interconversion Toolchain for the Embodied Data Pyramid
Peidong Liu, Yongce Liu, Songyan Guo +34
JoyAI-Sim is a toolchain that connects real robots, simulation, and human demonstrations to enable scalable evaluation and generation of robot training data using calibrated digita…
Towards a Unified Multidimensional Explainability Metric: Evaluating Trustworthiness in AI Models
Georgios Makridis, Georgios Fatouros, Athanasios Kiourtis +4
The paper proposes a framework that evaluates explainability methods like LIME and SHAP across models and datasets using fidelity, simplicity, and stability, and builds a knowledge…
Assessing AI in Introductory Physics Problem Solving
Amir Bralin, N. Sanjay Rebello
The paper evaluates OpenAI's o4-mini large language model on introductory physics problems from Halliday and Resnick, finding about 90% overall accuracy but lower performance on im…
The Model Knows Your Project, Not You: Measuring Recognition in LLMs with NameRank
Bojie Li, Noah Shi
The paper introduces NameRank, a metric that quantifies how well large language models can recognize specific people or tools from their internal weights without external retrieval…
VanillaBench: The Hidden Accuracy Cost of Adversarial Robustness
Niklas Bunzel
The paper presents VanillaBench, a benchmark that measures how much clean (vanilla) accuracy is lost when models are trained for adversarial robustness, revealing a larger accuracy…
Epistemic Stance Flexibility Probing: Measuring Prompt-Conditioned Register Shift in Large Language Models
Binwen Liu, Yilin Ren
The paper introduces ESFP, a benchmark that tests whether large language models can shift between neutral attribution and personal stance when prompted differently, and evaluates t…
SpurLens: Automatic Detection of Spurious Cues in Multimodal LLMs
Parsa Hosseini, Sumit Nawathe, Mazda Moayeri +2
The paper introduces SpurLens, an automated pipeline that uses GPT-4 and open-set object detectors to find spurious visual cues in multimodal large language models, showing that th…
Detecting Explanatory Insufficiency in Learned Representations: A Framework for Representational Vigilance
Jacques Raynal, Pierre Slangen, Elsa Raynal +1
The paper proposes VER, a conceptual framework for monitoring learned representations to detect unexplained residual structures that standard performance metrics miss, offering a d…
Trustworthy synthetic data for campaign decision support: strategy simulation fidelity and the PolicySynth framework
Tung Dang, The Hung Phung, Son Lam Nguyen +1
The paper presents PolicySynth, a synthetic data generation framework designed to preserve decision alignment for marketing campaign support, and introduces strategy simulation fid…
Characterising AI Models for Cataloguing
Miguel Arana-Catania, Neil Jefferies
The paper evaluates various AI models for automatically generating catalogue records for digitized collections, comparing implementations through qualitative and quantitative exper…
Where Experts Disagree, Models Fail: Detecting Implicit Legal Citations in French Court Decisions
Avrile Floro, Tamara Dhorasoo, Soline Pellez +1
The paper introduces a benchmark for detecting implicit citations of the French Civil Code in court decisions and shows that cases where legal experts disagree are especially hard…
Are DeepFakes Realistic Enough? Exploring Semantic Mismatch as a Novel Challenge
Sharayu Nilesh Deshmukh, Kailash A. Hambarde, Joana C. Costa +2
The paper introduces a new evaluation class for DeepFake detection that captures semantic inconsistencies between audio and video (RARV‑SMM) and tests how current models handle suc…
Evaluating Nonuniform Dependability Across Response Conditions: A Conditional Generalizability Framework Illustrated in Automated Essay Scoring
Yi Gui
The paper proposes a conditional generalizability framework to assess how reliable automated essay scoring systems are across different response conditions, using entropy-based str…
One search, two signals: results blend meaning (embedding similarity, so papers that never use your words still surface) with keyword matches on titles, abstracts and summaries. Free, no sign-in needed.