3 papers
cs.AI2026
Unbiased Prevalence Estimation with Multicalibrated LLMs
Fridolin Linder, Thomas Leeper, Daniel Haimovich +3
Estimating the prevalence of a category in a population using imperfect measurement devices (diagnostic tests, classifiers, or large language models) is fundamental to science, pub…
cs.LG2025
An Adaptive Approach for Infinitely Many-armed Bandits under Generalized Rotting Constraints
Jung-hun Kim, Milan Vojnovic, Se-Young Yun
In this study, we consider the infinitely many-armed bandit problems in a rested rotting setting, where the mean reward of an arm may decrease with each pull, while otherwise, it r…
cs.LG2025
What is the Alignment Objective of GRPO?
Milan Vojnovic, Se-Young Yun
In this note, we examine the aggregation of preferences achieved by the Group Policy Optimisation (GRPO) algorithm, a reinforcement learning method used to train advanced artificia…