4 papers
PoliTok-DE: A Multimodal Dataset of Political TikToks and Deletions From Germany
Tomas Ruiz, Andreas Nanz, Ursula Kristin Schmid +3
We present PoliTok-DE, a large-scale multimodal dataset (video, audio, images, text) of TikTok posts from two German elections: the 2024 Saxony state election and the 2025 German f…
Rethinking Ground Truth: A Case Study on Human Label Variation in MLLM Benchmarking
Tomas Ruiz, Tanalp Agustoslu, Carsten Schwemmer
Human Label Variation (HLV), i.e. systematic differences among annotators' judgments, remains underexplored in benchmarks despite rapid progress in large language model (LLM) devel…
Beyond Correctness: Evaluating and Improving LLM Feedback in Statistical Education
Niklas Ippisch, Markus Herklotz, Anna-Carolina Haensch +1
Large language models (LLMs) have been proposed as scalable tools to address the gap between the importance of individualized written feedback and the practical challenges of provi…
BoN Appetit Team at LeWiDi-2025: Best-of-N Test-time Scaling Can Not Stomach Annotation Disagreements (Yet)
Tomas Ruiz, Siyao Peng, Barbara Plank +1
Test-time scaling is a family of techniques to improve LLM outputs at inference time by performing extra computation. To the best of our knowledge, test-time scaling has been limit…