6 papers
Expert Evaluation and the Limits of Human Feedback in Mental Health AI Safety Testing
Kiana Jafari, Paul Ulrich Nikolaus Rust, Duncan Eddy +7
Learning from human feedback~(LHF) assumes that expert judgments, appropriately aggregated, yield valid ground truth for training and evaluating AI systems. We tested this assumpti…
The Doctor Will (Still) See You Now: On the Structural Limits of Agentic AI in Healthcare
Gabriela Aránguiz Dias, Kiana Jafari, Allie Griffith +4
Across healthcare, agentic artificial intelligence (AI) systems are increasingly promoted as capable of autonomous action, yet in practice they currently operate under near-total h…
An Adaptive Responsible AI Governance Framework for Decentralized Organizations
Kiana Jafari Meimandi, Anka Reuel, Gabriela Aranguiz-Dias +6
This paper examines the assessment challenges of Responsible AI (RAI) governance efforts in globally decentralized organizations through a case study collaboration between a leadin…
The Measurement Imbalance in Agentic AI Evaluation Undermines Industry Productivity Claims
Kiana Jafari Meimandi, Gabriela Aránguiz-Dias, Grace Ra Kim +3
As industry reports claim agentic AI systems deliver double-digit productivity gains and multi-trillion dollar economic potential, the validity of these claims has become critical…
More than Marketing? On the Information Value of AI Benchmarks for Practitioners
Amelia Hardy, Anka Reuel, Kiana Jafari Meimandi +6
Public AI benchmark results are widely broadcast by model developers as indicators of model quality within a growing and competitive market. However, these advertised scores do not…
Responsible AI in the Global Context: Maturity Model and Survey
Anka Reuel, Patrick Connolly, Kiana Jafari Meimandi +4
Responsible AI (RAI) has emerged as a major focus across industry, policymaking, and academia, aiming to mitigate the risks and maximize the benefits of AI, both on an organization…