2 papers
cs.LG2026
Establishing Construct Validity in LLM Capability Benchmarks Requires Nomological Networks
Timo Freiesleben
Recent work in machine learning increasingly attributes human-like capabilities such as reasoning or theory of mind to large language models (LLMs) on the basis of benchmark perfor…
cs.LG2025
The Benchmarking Epistemology: Construct Validity for Evaluating Machine Learning Models
Timo Freiesleben, Sebastian Zezulka
Predictive benchmarking, the evaluation of machine learning models based on predictive performance and competitive ranking, is a central epistemic practice in machine learning rese…