5 papers
Statistical Foundations of LLM-based A/B Testing: A Surrogacy Framework for Human Causal Inference
Joel Persson, MÃ¥rten Schultzberg, Sebastian Ankargren
Organizations and researchers show increasing interest in using large language models (LLMs) in place of human participants in A/B tests, in the hope of experimenting faster and at…
Logging Policy Design for Off-Policy Evaluation
Connor Douglas, Joel Persson, Foster Provost
Off-policy evaluation (OPE) estimates the value of a target treatment policy (e.g., a recommender system) using data collected by a different logging policy. It enables high-stakes…
EventChat: Implementation and user-centric evaluation of a large language model-driven conversational recommender system for exploring leisure events in an SME context
Hannes Kunstmann, Joseph Ollier, Joel Persson +1
Large language models (LLMs) present an enormous evolution in the strategic potential of conversational recommender systems (CRS). Yet to date, research has predominantly focused u…
Detecting and Mitigating Group Bias in Heterogeneous Treatment Effects
Joel Persson, Jurriën Bakker, Dennis Bohle +2
Heterogeneous treatment effects (HTEs) are increasingly estimated using machine learning models that produce highly personalized predictions of treatment effects. In practice, howe…
Auditing a Dutch Public Sector Risk Profiling Algorithm Using an Unsupervised Bias Detection Tool
Floris Holstege, Mackenzie Jorgensen, Kirtan Padh +4
Algorithms are increasingly used to automate or aid human decisions, yet recent research shows that these algorithms may exhibit bias across legally protected demographic groups. H…