6 citations · 7 across the 3 of their papers we have counts for
3 papers
cs.LG2026★ 1 cited
Establishing Construct Validity in LLM Capability Benchmarks Requires Nomological Networks
Timo Freiesleben
Recent work in machine learning increasingly attributes human-like capabilities such as reasoning or theory of mind to large language models (LLMs) on the basis of benchmark perfor…
cs.LG2025
The Benchmarking Epistemology: Construct Validity for Evaluating Machine Learning Models
Timo Freiesleben, Sebastian Zezulka
Predictive benchmarking, the evaluation of machine learning models based on predictive performance and competitive ranking, is a central epistemic practice in machine learning rese…
cs.AI2020★ 6 cited
The Intriguing Relation Between Counterfactual Explanations and Adversarial Examples
Timo Freiesleben
The same method that creates adversarial examples (AEs) to fool image-classifiers can be used to generate counterfactual explanations (CEs) that explain algorithmic decisions. This…