collaborators

5 papers

stat.ME2026

Statistical Foundations of LLM-based A/B Testing: A Surrogacy Framework for Human Causal Inference

Joel Persson, MÃ¥rten Schultzberg, Sebastian Ankargren

Organizations and researchers show increasing interest in using large language models (LLMs) in place of human participants in A/B tests, in the hope of experimenting faster and at…

stat.ML2026

Logging Policy Design for Off-Policy Evaluation

Connor Douglas, Joel Persson, Foster Provost

Off-policy evaluation (OPE) estimates the value of a target treatment policy (e.g., a recommender system) using data collected by a different logging policy. It enables high-stakes…

cs.IR2026

EventChat: Implementation and user-centric evaluation of a large language model-driven conversational recommender system for exploring leisure events in an SME context

Hannes Kunstmann, Joseph Ollier, Joel Persson +1

Large language models (LLMs) present an enormous evolution in the strategic potential of conversational recommender systems (CRS). Yet to date, research has predominantly focused u…

stat.ME2026

Detecting and Mitigating Group Bias in Heterogeneous Treatment Effects

Joel Persson, Jurriën Bakker, Dennis Bohle +2

Heterogeneous treatment effects (HTEs) are increasingly estimated using machine learning models that produce highly personalized predictions of treatment effects. In practice, howe…

cs.CY2026

Auditing a Dutch Public Sector Risk Profiling Algorithm Using an Unsupervised Bias Detection Tool

Floris Holstege, Mackenzie Jorgensen, Kirtan Padh +4

Algorithms are increasingly used to automate or aid human decisions, yet recent research shows that these algorithms may exhibit bias across legally protected demographic groups. H…