2 papers
cs.AI2025
AssistanceZero: Scalably Solving Assistance Games
Cassidy Laidlaw, Eli Bronstein, Timothy Guo +5
Assistance games are a promising alternative to reinforcement learning from human feedback (RLHF) for training AI assistants. Assistance games resolve key drawbacks of RLHF, such a…
cs.CY2024
GPAI Evaluations Standards Taskforce: Towards Effective AI Governance
Patricia Paskov, Lukas Berglund, Everett Smith +1
General-purpose AI evaluations have been proposed as a promising way of identifying and mitigating systemic risks posed by AI development and deployment. While GPAI evaluations pla…