activity
20242026
collaborators

6 papers

cs.AI2026

Expert Evaluation and the Limits of Human Feedback in Mental Health AI Safety Testing

Kiana Jafari, Paul Ulrich Nikolaus Rust, Duncan Eddy +7

Learning from human feedback~(LHF) assumes that expert judgments, appropriately aggregated, yield valid ground truth for training and evaluating AI systems. We tested this assumpti…

cs.CY2026

The Doctor Will (Still) See You Now: On the Structural Limits of Agentic AI in Healthcare

Gabriela Aránguiz Dias, Kiana Jafari, Allie Griffith +4

Across healthcare, agentic artificial intelligence (AI) systems are increasingly promoted as capable of autonomous action, yet in practice they currently operate under near-total h…

cs.CY2025

An Adaptive Responsible AI Governance Framework for Decentralized Organizations

Kiana Jafari Meimandi, Anka Reuel, Gabriela Aranguiz-Dias +6

This paper examines the assessment challenges of Responsible AI (RAI) governance efforts in globally decentralized organizations through a case study collaboration between a leadin…

cs.CY2025

The Measurement Imbalance in Agentic AI Evaluation Undermines Industry Productivity Claims

Kiana Jafari Meimandi, Gabriela Aránguiz-Dias, Grace Ra Kim +3

As industry reports claim agentic AI systems deliver double-digit productivity gains and multi-trillion dollar economic potential, the validity of these claims has become critical…

cs.AI2024

More than Marketing? On the Information Value of AI Benchmarks for Practitioners

Amelia Hardy, Anka Reuel, Kiana Jafari Meimandi +6

Public AI benchmark results are widely broadcast by model developers as indicators of model quality within a growing and competitive market. However, these advertised scores do not…

cs.CY2024

Responsible AI in the Global Context: Maturity Model and Survey

Anka Reuel, Patrick Connolly, Kiana Jafari Meimandi +4

Responsible AI (RAI) has emerged as a major focus across industry, policymaking, and academia, aiming to mitigate the risks and maximize the benefits of AI, both on an organization…