14 papers
LLM-SoccerArena: Benchmarking LLMs on Real-World Predictions in Sports
Jonas Schröder, Jonas Schweisthal, Oliver Müller +2
Large language models (LLMs) increasingly support decisions about uncertain future events, yet evaluating their ability to forecast real-world outcomes remains difficult. In partic…
Causal methods for LLM development and evaluation
Dennis Frauen, Marie Brockschmidt, Konstantin Hess +10
Large language model (LLM) development is currently driven by large-scale empirical iteration over data mixtures, reward models, routing strategies, and evaluation pipelines. Here,…
Adaptive Experimentation for Censored Survival Outcomes
Yuxin Wang, Dennis Frauen, Jonas Schweisthal +3
Adaptive experimentation enables efficient estimation of causal effects, but existing methods are not designed for survival data with censoring, where event times are only partiall…
Amortizing Causal Sensitivity Analysis via Prior Data-Fitted Networks
Emil Javurek, Dennis Frauen, Marie Brockschmidt +2
Causal sensitivity analysis aims to provide bounds for causal effect estimates in the presence of unobserved confounding. However, existing methods for causal sensitivity analysis…
Assessing the robustness of heterogeneous treatment effects in survival analysis under informative censoring
Yuxin Wang, Dennis Frauen, Jonas Schweisthal +2
Dropout is common in clinical studies, with up to half of patients leaving early due to side effects or other reasons. When dropout is informative (i.e., dependent on survival time…
A meta-analysis of the effect of generative AI on productivity and learning in programming
Sebastian Maier, Moritz Gunzenhäuser, Jonas Schweisthal +2
Generative artificial intelligence (GenAI) is increasingly used for programming, yet it remains unclear when and where GenAI tools lead to productivity gains. Evidence on the effec…