#model reliability

try —

4 papers match

cs.CV2026

Representation Trajectories Matters: Complementary Evidence for OOD Detection and Image Classification

Ignacio M. De la Jara, Cristian Rodriguez-Opazo, Hamed Damirchi +2

The paper investigates how the step‑by‑step changes in a vision model’s internal representations (representation trajectories) can be used to improve out‑of‑distribution detection…

#out-of-distribution detection#representation trajectories#image classification#intermediate layer analysis
cs.AI2026

Advancing Multimodal Judge Models through a Capability-Oriented Benchmark and MCTS-Driven Data Generation

Zeyu Chen, Huanjin Yao, Ziwang Zhao +1

The paper introduces a new benchmark, M-JudgeBench, to evaluate the judgment capabilities of multimodal large language models, and proposes a data generation method (Judge-MCTS) to…

#multimodal models#evaluation benchmarks#large language models#chain-of-thought
cs.CV2026

Where Not to Learn: Prior-Aligned Training with Subset-based Attribution Constraints for Reliable Decision-Making

Ruoyu Chen, Shangquan Sun, Xiaoqing Guo +8

The paper introduces a training approach that encodes human‑provided region priors and uses a subset‑selection attribution method to penalize models when their decision evidence fa…

#human prior alignment#attribution constraints#explainable AI#subset selection attribution
cs.LG2026

How Can Machine Learning Emulators Best Support Climate Science?

Luca Schmidt, Nina Effenberger, Vitus Benson +5

The paper examines how machine‑learning emulators can be designed and deployed to reduce the computational cost of physics‑based climate models, proposing a framework that emphasiz…

#climate modeling#machine learning emulators#computational efficiency#model reliability