3 papers
cs.LG2026
Position: Stop Evaluating AI with Human Tests, Develop Principled, AI-specific Tests instead
Tom Sühr, Florian E. Dorner, Olawale Salaudeen +2
Large Language Models (LLMs) have achieved remarkable results on a range of standardized tests originally designed to assess human cognitive and psychological traits, such as intel…
cs.LG2024
A Dynamic Model of Performative Human-ML Collaboration: Theory and Empirical Evidence
Tom Sühr, Samira Samadi, Chiara Farronato
Machine learning (ML) models are increasingly used in various applications, from recommendation systems in e-commerce to diagnosis prediction in healthcare. In this paper, we prese…
cs.LG2024
Online Decision Deferral under Budget Constraints
Mirabel Reid, Tom Sühr, Claire Vernade +1
Machine Learning (ML) models are increasingly used to support or substitute decision making. In applications where skilled experts are a limited resource, it is crucial to reduce t…