collaborators

5 papers

cs.AI2026

MentisOculi: Revealing the Limits of Reasoning with Mental Imagery

Jana Zeller, Thaddäus Wiedemer, Fanfei Li +6

Frontier models are transitioning from multimodal large language models (MLLMs) that merely ingest visual information to unified multimodal models (UMMs) capable of native interlea…

cs.CV2026

LAION-C: An Out-of-Distribution Benchmark for Web-Scale Vision Models

Fanfei Li, Thomas Klein, Wieland Brendel +2

Out-of-distribution (OOD) robustness is a desired property of computer vision models. Improving model robustness requires high-quality signals from robustness benchmarks to quantif…

cs.LG2026

Who Guards the Guardians? The Challenges of Evaluating Identifiability of Learned Representations

Shruti Joshi, Théo Saulus, Wieland Brendel +3

Identifiability in representation learning is commonly evaluated using standard metrics (e.g., MCC, DCI, R^2) on synthetic benchmarks with known ground-truth factors. These metrics…

cs.CV2026

Low-Pass Filtering Improves Behavioral Alignment of Vision Models

Max Wolff, Thomas Klein, Evgenia Rusak +2

Despite their impressive performance on computer vision benchmarks, Deep Neural Networks (DNNs) still fall short of adequately modeling human visual behavior, as measured by error…

q-bio.NC2025

Quantifying Uncertainty in Error Consistency: Towards Reliable Behavioral Comparison of Classifiers

Thomas Klein, Sascha Meyen, Wieland Brendel +2

Benchmarking models is a key factor for the rapid progress in machine learning (ML) research. Thus, further progress depends on improving benchmarking metrics. A standard metric to…