activity
20232026
most citedBias in the Loop: How Humans Evaluate AI-Generated Suggestions

2 citations · 3 across the 12 of their papers we have counts for

collaborators

25 papers

cs.AI2026

Automated reproducibility assessments in the social and behavioral sciences using large language models

Tobias Holtdirk, Pietro Marcolongo, Anna Steinberg Schulten +7

Reproducibility in the social and behavioral sciences is typically evaluated by independent researchers who reanalyze the original data to assess whether the published findings can…

cs.HC2026

AI Conversational Interviewing: Scaling Up Semi-Structured and In-depth Interviews

Alexander Wuttke, Max Melchior Lang, Christopher Klamm +2

Public opinion research has long faced a trade-off between depth and scale: standardized surveys enable large-scale measurement but restrict respondents to researcher-defined categ…

cs.LG2026

How Hard is it to Rig a Benchmark? A Social Choice Analysis of Leaderboard Robustness

Polina Gordienko, Georg Schollmeyer, Frauke Kreuter +1

Multi-task benchmarks have become a central pillar of machine learning research, yet their growing influence has incentivised benchmark gaming -- strategic actions taken to improve…

stat.ME2026

From Ground Truth to Measurement: A Statistical Framework for Human Labeling

Robert Chew, Stephanie Eckman, Christoph Kern +1

Supervised machine learning assumes that labeled data provide accurate measurements of the concepts models are meant to learn. Yet in practice, human labeling introduces systematic…

cs.CY2026

Reading Between the Tokens: Improving Preference Predictions through Mechanistic Forecasting

Sarah Ball, Simeon Allmendinger, Frauke Kreuter +1

Large language models are increasingly used to predict human preferences in both scientific and business endeavors, yet current approaches rely exclusively on analyzing model outpu…

cs.CL2026

Moral Lenses, Political Coordinates: Towards Ideological Positioning of Morally Conditioned LLMs

Chenchen Yuan, Bolei Ma, Zheyu Zhang +3

While recent research has systematically documented political orientation in large language models (LLMs), existing evaluations rely primarily on direct probing or demographic pers…