Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
LLM-SoccerArena: Benchmarking LLMs on Real-World Predictions in Sports
Jonas Schröder, Jonas Schweisthal, Oliver Müller +2
Large language models (LLMs) increasingly support decisions about uncertain future events, yet evaluating their ability to forecast real-world outcomes remains difficult. In partic…
cs.AI2026
Automated reproducibility assessments in the social and behavioral sciences using large language models
Tobias Holtdirk, Pietro Marcolongo, Anna Steinberg Schulten +7
Reproducibility in the social and behavioral sciences is typically evaluated by independent researchers who reanalyze the original data to assess whether the published findings can…