4 papers
Principles and Guidelines for Randomized Controlled Trials in AI Evaluation
Christopher Kelly, Angelica Chowdhury, Alexandra Campili +5
This work establishes a framework for standardizing AI evaluation RCTs (sometimes called human uplift studies). Drawing on established practices from disciplines with established R…
The Case for ESM3 as a General-Purpose AI Model with Systemic Risk Under the EU AI Act
Taro Qureshi, Jacob Griffith, Koen Holtman +3
Due to ambiguity in the wording of the EU AI Act, we examine the question of to what extent frontier biological foundation models such as ESM3 are subject to obligations for genera…
Deprecating Benchmarks: Criteria and Framework
Ayrton San Joaquin, Rokas Gipiškis, Leon Staufer +1
As frontier artificial intelligence (AI) models rapidly advance, benchmarks are integral to comparing different models and measuring their progress in different task-specific domai…
PreCare: Designing AI Assistants for Advance Care Planning (ACP) to Enhance Personal Value Exploration, Patient Knowledge, and Decisional Confidence
Yu Lun Hsu, Yun-Rung Chou, Chiao-Ju Chang +9
Advance Care Planning (ACP) allows individuals to specify their preferred end-of-life life-sustaining treatments before they become incapacitated by injury or terminal illness (e.g…