collaborators

8 papers

cs.AI2026

TriQua: Reconciling Granularity and Context in Factuality Evaluation

Jin Liu, Steffen Thoma, Achim Rettinger

The "decompose-then-verify" paradigm for LLM factuality evaluation faces a fundamental trade-off: atomic facts, i.e., one sentence conveying one unit of information, often omit ess…

cs.CL2026

The Unsampled Truth: Psychometrics in SLMs Measure Prompt Artifacts, Not Psychological Constructs

Nils Schwager, Christoph Hau, Simon Münker +1

When prompting SLMs for psychometric assessments, researchers assume the outputs reflect semantic reasoning. We evaluate this premise across 13 open-weights models (0.6B to 14B par…

cs.CL2026

LLM Agents Predict Social Media Reactions but Do Not Outperform Text Classifiers: Benchmarking Simulation Accuracy Using 120K+ Personas of 1511 Humans

Ljubisa Bojic, Alexander Felfernig, Bojana Dinic +4

Social media platforms mediate how billions form opinions and engage with public discourse. As autonomous AI agents increasingly participate in these spaces, understanding their be…

cs.CL2026

Towards Simulating Social Media Users with LLMs: Evaluating the Operational Validity of Conditioned Comment Prediction

Nils Schwager, Simon Münker, Alistair Plum +1

The transition of Large Language Models (LLMs) from exploratory tools to active "silicon subjects" in social science lacks extensive validation of operational validity. This study…

cs.CL2026

Next Reply Prediction X Dataset: Linguistic Discrepancies in Naively Generated Content

Simon Münker, Nils Schwager, Kai Kugler +2

The increasing use of Large Language Models (LLMs) as proxies for human participants in social science research presents a promising, yet methodologically risky, paradigm shift. Wh…

cs.CL2025

Identity-Aware Large Language Models require Cultural Reasoning

Alistair Plum, Anne-Marie Lutgen, Christoph Purschke +1

Large language models have become the latest trend in natural language processing, heavily featuring in the digital tools we use every day. However, their replies often reflect a n…